paper-with-me

홈 › Papers

VUDG: A Dataset for Video Understanding Domain Generalization

2025-05-30 · Ziyi Wang, Zhi Gao, Boxuan Yu, Zirui Dai, Yuxiang Song, Qingyuan Lu, Jin Chen, Xinxiao wu

Video understanding has made remarkable progress in recent years, largely driven by advances in deep models and the availability of large-scale annotated datasets. However, existing works typically ignore the inherent domain shifts encountered in real-world video applications, leaving domain generalization (DG) in video understanding underexplored. Hence, we propose Video Understanding Domain Generalization (VUDG), a novel dataset designed specifically for evaluating the DG performance in video understanding. VUDG contains videos from 11 distinct domains that cover three types of domain shifts, and maintains semantic similarity across different domains to ensure fair and meaningful evaluation. We propose a multi-expert progressive annotation framework to annotate each video with both multiple-choice and open-ended question-answer pairs. Extensive experiments on 9 representative large video-language models (LVLMs) and several traditional video question answering methods show that most models (including state-of-the-art LVLMs) suffer performance degradation under domain shifts. These results highlight the challenges posed by VUDG and the difference in the robustness of current models to data distribution shifts. We believe VUDG provides a valuable resource for prompting future research in domain generalization video understanding.

📄 PDF Abstract BibTeX arXiv:2505.24346

Code (0)

등록된 구현이 없습니다.

Tasks

Domain GeneralizationMultiple-choiceQuestion AnsweringSemantic SimilaritySemantic Textual SimilarityVideo Question AnsweringVideo Understanding

Similar Papers 제목 키워드 기반

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

2026-07-16 · Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong 외 arxiv

Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in se…

Computational Efficiency

Enhancing Video Transformers for Action Understanding with VLM-aided Training

2024-03-24 · Hui Lu, Hu Jian, Ronald Poppe, Albert Ali Salah

Owing to their ability to extract relevant spatio-temporal video embeddings, Vision Transformers (ViTs) are currently the best performing models in video action understanding. However, their generalization over domains o…

Action ClassificationAction RecognitionAction Understanding

How Severe is Benchmark-Sensitivity in Video Self-Supervised Learning?

2022-03-27 · Fida Mohammad Thoker, Hazel Doughty, Piyush Bagad, Cees Snoek

Despite the recent success of video self-supervised learning models, there is much still to be understood about their generalization capability. In this paper, we investigate how sensitive video self-supervised learning …

Self-Supervised LearningSensitivityVideo Understanding

ViP: Video Platform for PyTorch

2019-10-07 · Madan Ravi Ganesh, Eric Hofesmann, Nathan Louis, Jason Corso

This work presents the Video Platform for PyTorch (ViP), a deep learning-based framework designed to handle and extend to any problem domain based on videos. ViP supports (1) a single unified interface applicable to all …

BenchmarkingVideo Understanding

Unified Video Dense Prediction from Disjoint Data

2026-07-23 · Yihong Sun, Seoung Wug Oh, Jiahui Huang, Bharath Hariharan 외 arxiv

Scene understanding requires simultaneous prediction about geometry, appearance, and semantics. However, existing task-specific annotations are fragmented across incompatible, domain-specific datasets. Current unified sy…

Semantic SegmentationScene Understanding