paper-with-me

홈 › Papers

OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains

2026-06-12 · Xinyue Cai, Chaoyou Fu, Yi-Fan Zhang, Ran He, Caifeng Shan arxiv

Current automated pipelines for audio-visual Question Answering (QA) generally adopt a ``video-caption-QA'' paradigm. However, these methods typically segment videos into short clips and generate separate descriptions for audio and visual modalities. This decoupled processing severs inherent associations between sounds and their visual sources, while independent clip processing often causes inconsistent descriptions of the same entity across segments. Furthermore, coupling long-text comprehension and QA synthesis into a single step often restricts models to localized events, yielding questions lacking long-term temporal connections and deep cross-modal reasoning. To address these issues, we propose an automated data engine featuring two mechanisms: (1) \textbf{Entity-Anchored Video Scripting} transforms videos into structured scripts, comprising summaries, main entity lists, and segment-wise audio-visual descriptions. The entity list serves as a global prior to ensure cross-segment referential consistency and reconstruct audio-visual associations. (2) \textbf{Clue-Guided QA Generation} prompts models to first mine cross-segment, multimodal clues from the script, and subsequently generate QA pairs based on these high-value clues. Leveraging this pipeline, we construct the instruction-tuning dataset \textbf{OmniVideo-100K} and a human-verified test set, \textbf{OmniVideo-Test}. Fine-tuning VITA-1.5, Qwen2.5-Omni-7B and Qwen3-Omni-30B on OmniVideo-100K yields performance gains of up to 20.59% on OmniVideo-Test, demonstrating strong generalization (up to 12.64% improvements) across established benchmarks like Daily-Omni and JointAVBench.

📄 PDF Abstract BibTeX arXiv:2606.14702

Code (0)

등록된 구현이 없습니다.

Tasks

Audio-visual Question AnsweringVisual Reasoning

Similar Papers 제목 키워드 기반

OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention

2026-02-05 · Zhangquan Chen, Jiale Tao, Ruihuang Li, Yihao Hu 외 arxiv

While humans perceive the world through diverse modalities that operate synergistically to support a holistic understanding of their surroundings, existing omnivideo models still face substantial challenges on audio-visu…

Self-Supervised LearningContrastive LearningVisual Reasoning

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

2025-10-12 · Caorui Li, Yu Chen, Yiyan Ji, Jin Xu 외 arxiv

Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities…

Causal InferenceVisual Reasoning

Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning

2026-05-27 · Ke Xu, Yuhao Wang, Ziyang Cheng, Hongcheng Liu 외 arxiv

Multi-hop audio-visual reasoning remains challenging for Omni-LLMs, as relevant evidence is often sparse, temporally dispersed, and distributed across both audio and visual streams. Existing benchmarks provide limited in…

Visual Reasoning

OmniRAG-Agent: Agentic Omnimodal Reasoning for Low-Resource Long Audio-Video Question Answering

2026-02-03 · Yifan Zhu, Xinyu Mu, Tao Feng, Zhonghong Ou 외 arxiv

Long-horizon omnimodal question answering answers questions by reasoning over text, images, audio, and video. Despite recent progress on OmniLLMs, low-resource long audio-video QA still suffers from costly dense encoding…

Video Question Answering

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers

2026-04-17 · Edson Araujo, Saurabhchand Bhati, M. Jehanzeb Mirza, Brian Kingsbury 외 arxiv

Recent advances in reasoning models have shown remarkable progress in text-based domains, but transferring those capabilities to multimodal settings, e.g., to allow reasoning over audio-visual data, still remains a chall…

Reinforcement LearningMultimodal ReasoningVisual Reasoning