paper-with-me

홈 › Papers

Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning

2026-05-27 · Ke Xu, Yuhao Wang, Ziyang Cheng, Hongcheng Liu, Yanfeng Wang, Yu Wang arxiv

Multi-hop audio-visual reasoning remains challenging for Omni-LLMs, as relevant evidence is often sparse, temporally dispersed, and distributed across both audio and visual streams. Existing benchmarks provide limited investigation of this setting, typically involving only a limited number of modalities, relevant temporal segments, or reasoning steps. In this work, we introduce MOV-Bench, a benchmark containing 519 carefully curated questions that require multi-hop reasoning over temporally dispersed audio-visual evidence. Evaluations on MOV-Bench reveal that current Omni-LLMs still struggle with multi-hop cross-modal reasoning. To address this challenge, we further propose AOP-Agent, an efficient agentic framework built on open-source Omni-LLMs for active omni-modal perception. By combining a hierarchical omni-modal memory with a collaborative observe-reflect-replan loop, AOP-Agent enables open-source Omni-LLMs to perform active perception without additional training or proprietary models. Experiments on MOV-Bench and OmniVideoBench demonstrate that AOP-Agent consistently improves reasoning performance, with particularly notable gains on long videos and reasoning-intensive questions.

📄 PDF Abstract BibTeX arXiv:2605.28192

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

Native Active Perception as Reasoning for Omni-Modal Understanding

2026-06-17 · Zhenghao Xing, Ruiyang Xu, Yuxuan Wang, Jinzheng He 외 arxiv

Passive models for long video understanding typically rely on a "watch-it-all" paradigm, processing frames uniformly regardless of query difficulty, causing computational cost to grow with video duration. Although intera…

Reinforcement Learning

Active Perception Agent for Omnimodal Audio-Video Understanding

2025-12-29 · Keda Tao, Wenjie Du, Bohan Yu, Weiqiang Wang 외 arxiv

Omnimodal large language models have made significant strides in unifying audio and visual modalities; however, they often face challenges in fine-grained cross-modal understanding and have difficulty with multimodal ali…

Response Generation

Omni Interaction Agent Technical Report

2026-09-08 · Orantqing, Shengpeng Ji, Junlong Tong, Jialong Zuo 외 arxiv

In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and agentic capabilities within a single framework. In contrast to turn-based conventional paradigms, Gander contin…

OmniGAIA: Towards Native Omni-Modal AI Agents

2026-02-26 · Xiaoxi Li, Wenxiang Jiao, Jiarui Jin, Haoxuan Li 외 arxiv

Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to interact with the world. However, current multi-modal LLMs are primarily …

OmniRAG-Agent: Agentic Omnimodal Reasoning for Low-Resource Long Audio-Video Question Answering

2026-02-03 · Yifan Zhu, Xinyu Mu, Tao Feng, Zhonghong Ou 외 arxiv

Long-horizon omnimodal question answering answers questions by reasoning over text, images, audio, and video. Despite recent progress on OmniLLMs, low-resource long audio-video QA still suffers from costly dense encoding…

Video Question Answering