paper-with-me

홈 › Papers

WildDepth: A Multimodal Dataset for 3D Wildlife Perception and Depth Estimation

2026-03-17 · Muhammad Aamir, Naoya Muramatsu, Sangyun Shin, Matthew Wijers, Jia-Xing Zhong, Xinyu Hou, Amir Patel, Andrew Loveridge, Andrew Markham arxiv

Depth estimation and 3D reconstruction have been extensively studied as core topics in computer vision. Starting from rigid objects with relatively simple geometric shapes, such as vehicles, the research has expanded to address general objects, including challenging deformable objects, such as humans and animals. However, for the animal, in particular, the majority of existing models are trained based on datasets without metric scale, which can help validate image-only models. To address this limitation, we present WildDepth, a multimodal dataset and benchmark suite for depth estimation, behavior detection, and 3D reconstruction from diverse categories of animals ranging from domestic to wild environments with synchronized RGB and LiDAR. Experimental results show that the use of multi-modal data improves depth reliability by up to 10% RMSE, while RGB-LiDAR fusion enhances 3D reconstruction fidelity by 12% in Chamfer distance. By releasing WildDepth and its benchmarks, we aim to foster robust multimodal perception systems that generalize across domains.

📄 PDF Abstract BibTeX arXiv:2603.16816

Code (0)

등록된 구현이 없습니다.

Tasks

3D ReconstructionDepth Estimation

Similar Papers 제목 키워드 기반

WildBox: A Dataset and Benchmark for Aerial Monocular 3D Detection of African Savanna Wildlife

2026-06-19 · Vandita Shukla, Kilian Meier, Lucie Laporte-Devylder, Camille Rondeau Saint-Jean 외 arxiv

We introduce WildBox, a dataset and benchmark for monocular 3D detection of wildlife from drone video, comprising 237,505 3D bounding box annotations across seven African savanna species grouped into six benchmark classe…

Benchmark on Monocular Metric Depth Estimation in Wildlife Setting

2025-10-06 · Niccolò Niccoli, Lorenzo Seidenari, Ilaria Greco, Francesco Rovero arxiv

Camera traps are widely used for wildlife monitoring, but extracting accurate distance measurements from monocular images remains challenging due to the lack of depth information. While monocular depth estimation (MDE) m…

Monocular Depth EstimationComputational Efficiency

SmartWilds: Multimodal Wildlife Monitoring Dataset

2025-09-23 · Jenna Kline, Anirudh Potlapally, Bharath Pillai, Tanishka Wani 외 arxiv

We present the first release of SmartWilds, a multimodal wildlife monitoring dataset. SmartWilds is a synchronized collection of drone imagery, camera trap photographs and videos, and bioacoustic recordings collected dur…

Wildlife Product Trading in Online Social Networks: A Case Study on Ivory-Related Product Sales Promotion Posts

2024-09-25 · Guanyi Mou, Yun Yue, Kyumin Lee, Ziming Zhang

Wildlife trafficking (WLT) has emerged as a global issue, with traffickers expanding their operations from offline to online platforms, utilizing e-commerce websites and social networks to enhance their illicit trade. Th…

Perception Tokens Enhance Visual Reasoning in Multimodal Language Models

2024-12-04 · CVPR 2025 1 · Mahtab Bigverdi, Zelun Luo, Cheng-Yu Hsieh, Ethan Shen 외

Multimodal language models (MLMs) still face challenges in fundamental visual perception tasks where specialized models excel. Tasks requiring reasoning about 3D structures benefit from depth estimation, and reasoning ab…

Depth Estimationobject-detectionObject DetectionVisual Reasoning