paper-with-me

홈 › Papers

Web-Scale Collection of Video Data for 4D Animal Reconstruction

2025-11-03 · Brian Nlong Zhao, Jiajun Wu, Shangzhe Wu arxiv

Computer vision for animals holds great promise for wildlife research but often depends on large-scale data, while existing collection methods rely on controlled capture setups. Recent data-driven approaches show the potential of single-view, non-invasive analysis, yet current animal video datasets are limited--offering as few as 2.4K 15-frame clips and lacking key processing for animal-centric 3D/4D tasks. We introduce an automated pipeline that mines YouTube videos and processes them into object-centric clips, along with auxiliary annotations valuable for downstream tasks like pose estimation, tracking, and 3D/4D reconstruction. Using this pipeline, we amass 30K videos (2M frames)--an order of magnitude more than prior works. To demonstrate its utility, we focus on the 4D quadruped animal reconstruction task. To support this task, we present Animal-in-Motion (AiM), a benchmark of 230 manually filtered sequences with 11K frames showcasing clean, diverse animal motions. We evaluate state-of-the-art model-based and model-free methods on Animal-in-Motion, finding that 2D metrics favor the former despite unrealistic 3D shapes, while the latter yields more natural reconstructions but scores lower--revealing a gap in current evaluation. To address this, we enhance a recent model-free approach with sequence-level optimization, establishing the first 4D animal reconstruction baseline. Together, our pipeline, benchmark, and baseline aim to advance large-scale, markerless 4D animal reconstruction and related tasks from in-the-wild videos. Code and datasets are available at https://github.com/briannlongzhao/Animal-in-Motion.

📄 PDF Abstract BibTeX arXiv:2511.01169

Code (0)

등록된 구현이 없습니다.

Tasks

Pose Estimation

Similar Papers 제목 키워드 기반

Ponymation: Learning 3D Animal Motions from Unlabeled Online Videos

2023-12-21 · Keqiang Sun, Dor Litvak, Yunzhi Zhang, Hongsheng Li 외

We introduce Ponymation, a new method for learning a generative model of articulated 3D animal motions from raw, unlabeled online videos. Unlike existing approaches for motion synthesis, our model does not require any po…

Motion Synthesis

Kirin: Animal Motion Generation from In-the-Wild Video

2026-09-01 · Brian Nlong Zhao, Zhuoyang Pan, James M. Rehg, Jiajun Wu 외 hf

Understanding animal motion is fundamental to modeling animal behavior and biomechanics, yet progress in this area lags far behind human motion research due to the scarcity of high-quality motion data. While human motion…

Reinforcement Learning from Wild Animal Videos

2024-12-05 · Elliot Chane-Sane, Constant Roux, Olivier Stasse, Nicolas Mansard

We propose to learn legged robot locomotion skills by watching thousands of wild animal videos from the internet, such as those featured in nature documentaries. Indeed, such videos offer a rich and diverse collection of…

reinforcement-learningReinforcement Learning

Unleashing Infinite Motion: Scaling Expressive Quadrupedal Motion via Generative Video Priors

2026-06-26 · Youzhi Liu, Li Gao, Yifei Qian, Liu Liu 외 arxiv

Quadruped robots have achieved remarkable locomotion, yet their behavioral repertoire remains confined to a few gaits--far from the expressive, companion-like presence long envisioned for them. Attempts to import the hum…

Online Adaptation for Consistent Mesh Reconstruction in the Wild

2020-12-06 · NeurIPS 2020 12 · Xueting Li, Sifei Liu, Shalini De Mello, Kihwan Kim 외

This paper presents an algorithm to reconstruct temporally consistent 3D meshes of deformable object instances from videos in the wild. Without requiring annotations of 3D mesh, 2D keypoints, or camera pose for each vide…

3D Reconstruction