paper-with-me

Papers

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search

2026-05-09 · Tao Yu, yiming ding, Shenghua Chai, Minghui Zhang, Zhongtian Luo, Xinming Wang, Xinlong Chen, Zhaolu Kang, Junhao Gong, Yuxuan Zhou, Haopeng Jin, Zhiqing Cui, Jiabing Yang, YiFan Zhang, Hongzhu Yi, Zheqi He, Xi Yang, Yan Huang, Liang Wang arxiv

Current omni-modal benchmarks mainly evaluate models under settings where multiple modalities are provided simultaneously, while the ability to start from audio alone and actively search for cross-modal evidence remains underexplored. In this paper, we introduce \textbf{Omni-DeepSearch}, a benchmark for audio-driven omni-modal deep search. Given one or more audio clips and a related question, models must infer useful clues from audio, invoke text, image, and video search tools, and perform multi-hop reasoning to produce a short, objective, and verifiable answer. Omni-DeepSearch contains 640 samples across 15 fine-grained categories, covering four retrieval target modalities and four audio content types. A multi-stage filtering pipeline ensures audio dependence, retrieval necessity, visual modality necessity, and answer uniqueness. Experiments on recent closed-source and open-source omni-modal models show that this task remains highly challenging: the strongest evaluated model, Gemini-3-Pro, achieves only 43.44\% average accuracy. Further analyses illustrate key bottlenecks in audio entity inference, query formulation, tool-use reliability, multi-hop retrieval, and cross-modal verification. These results highlight audio-driven omni-modal deep search as an important and underexplored direction for future multimodal agents.

📄 PDF Abstract BibTeX arXiv:2605.08762

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

2026-07-28 · Haoyang Huang, Wenjie Huang, Tianqi Xu, Hongyaoxing Gu 외 arxiv

Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs. Existing compressi…

Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection

2026-05-02 · Tianxiao Li, Zhenglin Huang, Haiquan Wen, Yiwei He 외 arxiv

Multimodal deepfakes are proliferating on social media and threaten authenticity, information integrity, and digital forensics. Existing benchmarks are constrained by their single-modality scope, simplified manipulations…

DeepFake Detection

Omni-o3: Deep Nested Omnimodal Deduction for Deliberative Audio-Visual Reasoning

2026-04-27 · Zhicheng Zhang, Wentao Gu, Weicheng Wang, Yongjie Zhu 외 arxiv

Omnimodal understanding entails a massive, highly redundant search space of cross-modal interactions, demanding focused and deliberative reasoning. Current reasoning paradigms rely on either sequential step-by-step gener…

Reinforcement LearningVisual Reasoning

Omni-Interactive Universal Embedder

2026-08-27 · Wei-Yao Wang, Kazuya Tateishi, Shuyang Cui, Christian Simon 외 arxiv

Multimodal representation learning has been shifting from traditional two-tower architectures to large language model (LLM)-based embedders due to their strong instruction-following capabilities. Despite this progress, e…

Representation Learning

OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation

2025-06-23 · Qijun Gan, Ruizi Yang, Jianke Zhu, Shaofei Xue 외

Significant progress has been made in audio-driven human animation, while most existing methods focus mainly on facial movements, limiting their ability to create full-body animations with natural synchronization and flu…

Human AnimationVideo Generation