paper-with-me

홈 › Papers

Towards Understanding Camera Motions in Any Video

2025-04-21 · Zhiqiu Lin, Siyuan Cen, Daniel Jiang, Jay Karhade, Hewei Wang, Chancharik Mitra, Tiffany Ling, Yuhan Huang, Sifan Liu, Mingyu Chen, Rushikesh Zawar, Xue Bai, Yilun Du, Chuang Gan, Deva Ramanan

We introduce CameraBench, a large-scale dataset and benchmark designed to assess and improve camera motion understanding. CameraBench consists of ~3,000 diverse internet videos, annotated by experts through a rigorous multi-stage quality control process. One of our contributions is a taxonomy of camera motion primitives, designed in collaboration with cinematographers. We find, for example, that some motions like "follow" (or tracking) require understanding scene content like moving subjects. We conduct a large-scale human study to quantify human annotation performance, revealing that domain expertise and tutorial-based training can significantly enhance accuracy. For example, a novice may confuse zoom-in (a change of intrinsics) with translating forward (a change of extrinsics), but can be trained to differentiate the two. Using CameraBench, we evaluate Structure-from-Motion (SfM) and Video-Language Models (VLMs), finding that SfM models struggle to capture semantic primitives that depend on scene content, while VLMs struggle to capture geometric primitives that require precise estimation of trajectories. We then fine-tune a generative VLM on CameraBench to achieve the best of both worlds and showcase its applications, including motion-augmented captioning, video question answering, and video-text retrieval. We hope our taxonomy, benchmark, and tutorials will drive future efforts towards the ultimate goal of understanding camera motions in any video.

📄 PDF Abstract BibTeX arXiv:2504.15376

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringText RetrievalVideo Question AnsweringVideo-Text Retrieval

Similar Papers 제목 키워드 기반

Learning to Film From Professional Human Motion Videos

2019-06-01 · CVPR 2019 6 · Chong Huang, Chuan-En Lin, Zhenyu Yang, Yan Kong 외

We investigate the problem of 6 degrees of freedom (DOF) camera planning for filming professional human motion videos using a camera drone. Existing methods either plan motions for only a pan-tilt-zoom (PTZ) camera, or …

MotionMaster: Training-free Camera Motion Transfer For Video Generation

2024-04-24 · Teng Hu, Jiangning Zhang, Ran Yi, Yating Wang 외

The emergence of diffusion models has greatly propelled the progress in image and video generation. Recently, some efforts have been made in controllable video generation, including text-to-video generation and video mot…

DisentanglementMotion DisentanglementText-to-Video GenerationVideo Generation

VividCam: Learning Unconventional Camera Motions from Virtual Synthetic Videos

2025-10-28 · Qiucheng Wu, Handong Zhao, Zhixin Shu, Jing Shi 외 arxiv

Although recent text-to-video generative models are getting more capable of following external camera controls, imposed by either text descriptions or camera trajectories, they still struggle to generalize to unconventio…

MotionFlow:Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation

2025-09-25 · Guojun Lei, Chi Wang, Yikai Wang, Hong Li 외 arxiv

Generating videos guided by camera trajectories poses significant challenges in achieving consistency and generalizability, particularly when both camera and object motions are present. Existing approaches often attempt …

Video Generation

ViMo: Generating Motions from Casual Videos

2024-08-13 · Liangdong Qiu, Chengxing Yu, Yanran Li, Zhao Wang 외

Although humans have the innate ability to imagine multiple possible actions from videos, it remains an extraordinary challenge for computers due to the intricate camera movements and montages. Most existing motion gener…

Motion Generation