paper-with-me

홈 › Papers

Video Understanding by Design: How Datasets Shape Video Models

2025-09-11 · Lei Wang, Syuan-Hao Li, Piotr Koniusz, Yongsheng Gao arxiv

Research in video understanding has advanced rapidly, driven by increasingly diverse datasets and more powerful model architectures. While existing surveys typically organize progress by tasks, benchmarks, or model families, they provide limited insight into why particular architectures emerged and succeeded. In this survey, we argue that the evolution of video understanding is fundamentally shaped by dataset structure. We present a dataset-centric perspective that connects dataset structure, inductive biases, and architectural design within a unified framework. We show that different datasets require models to capture specific invariances and capabilities, such as robustness to viewpoint changes, sensitivity to temporal ordering, reasoning over long-range dependencies, relational interactions, and cross-modal alignment. These requirements naturally give rise to inductive biases, i.e., architectural assumptions that favor particular patterns of reasoning and generalization. From this perspective, milestone architectures, including two-stream networks, 3D CNNs, temporal models, transformers, graph-based methods, and multimodal foundation models, can be understood as architectural responses to the challenges posed by evolving datasets. Building on this framework, we systematically analyze how dataset characteristics have shaped architectural innovation across video understanding tasks and discuss the representational biases induced by different data regimes. By unifying datasets, inductive biases, and architectures into a coherent perspective, this survey offers both a retrospective explanation of the field's evolution and a forward-looking roadmap toward general-purpose video understanding systems. Code and dynamic video visualizations of dataset-induced biases are available at https://time.griffith.edu.au/paper-sites/video-understanding/.

📄 PDF Abstract BibTeX arXiv:2509.09151

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos

2024-11-18 · Yunong Liu, Cristobal Eyzaguirre, Manling Li, Shubh Khanna 외

Shape assembly is a ubiquitous task in daily life, integral for constructing complex 3D structures like IKEA furniture. While significant progress has been made in developing autonomous agents for shape assembly, existin…

Pose EstimationSemantic SegmentationVideo Object SegmentationVideo Semantic Segmentation

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

2025-01-22 · Boqiang Zhang, Kehan Li, Zesen Cheng, Zhiqiang Hu 외

In this paper, we propose VideoLLaMA3, a more advanced multimodal foundation model for image and video understanding. The core design philosophy of VideoLLaMA3 is vision-centric. The meaning of "vision-centric" is two-fo…

PhilosophyVideo Question AnsweringVideo Understanding

3D Shape Temporal Aggregation for Video-Based Clothing-Change Person Re-Identication

2023-03-09 · Asian Conference on Computer Vision 2023 3 · Ke Han, Shaogang Gong, Yan Huang, Liang Wang 외

3D shape of human body can be both discriminative and clothing-independent information in video-based clothing-change person re-identification (Re-ID). However, existing Re-ID methods usually generate 3D body shapes with…

3D Shape GenerationPerson Re-Identification

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding

2025-05-29 · David Ma, Huaqing Yuan, Xingjian Wang, Qianbo Zang 외

Although long-video understanding demands that models capture hierarchical temporal information -- from clip (seconds) and shot (tens of seconds) to event (minutes) and story (hours) -- existing benchmarks either neglect…

AvgVideo Understanding

Learning Transferable Temporal Primitives for Video Reasoning via Synthetic Videos

2026-03-18 · Songtao Jiang, Sibo Song, Chenyi Zhou, Yuan Wang 외 arxiv

The transition from image to video understanding requires vision-language models (VLMs) to shift from recognizing static patterns to reasoning over temporal dynamics such as motion trajectories, speed changes, and state …

Video Generation