paper-with-me

홈 › Papers

Native Active Perception as Reasoning for Omni-Modal Understanding

2026-06-17 · Zhenghao Xing, Ruiyang Xu, Yuxuan Wang, Jinzheng He, Ziyang Ma, Qize Yang, Yunfei Chu, Jin Xu, Junyang Lin, Chi-Wing Fu, Pheng-Ann Heng arxiv

Passive models for long video understanding typically rely on a "watch-it-all" paradigm, processing frames uniformly regardless of query difficulty, causing computational cost to grow with video duration. Although interactive frameworks have emerged, they often rely on global pre-scanning, and their context cost still scales with video length. We propose OmniAgent, the first native omni-modal agent that formulates video understanding as a POMDP-based iterative Observation-Thought-Action cycle. OmniAgent executes on-demand actions to selectively distill audio-visual cues into a persistent textual memory, effectively decoupling reasoning complexity from raw video duration. To operationalize this, we introduce (1) Agentic Supervised Fine-Tuning to bootstrap native active perception via best-of-N trajectory synthesis with dual-stage quality control, and (2) Agentic Reinforcement Learning with TAURA (Turn-aware Adaptive Uncertainty Rescaled Advantage), which leverages turn-level entropy to steer credit assignment toward pivotal discovery turns. Crucially, OmniAgent exhibits positive test-time scaling, where performance improves as the number of reasoning turns increases, validating the efficacy of active perception. Empirical results across ten benchmarks (e.g., VideoMME, LVBench) demonstrate that OmniAgent achieves state-of-the-art performance among open-source models. Notably, on LVBench, our 7B agent outperforms the 10$\times$ larger Qwen2.5-VL-72B (50.5% vs. 47.3%).

📄 PDF Abstract BibTeX arXiv:2606.19341

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

OmniGAIA: Towards Native Omni-Modal AI Agents

2026-02-26 · Xiaoxi Li, Wenxiang Jiao, Jiarui Jin, Haoxuan Li 외 arxiv

Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to interact with the world. However, current multi-modal LLMs are primarily …

Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning

2026-05-27 · Ke Xu, Yuhao Wang, Ziyang Cheng, Hongcheng Liu 외 arxiv

Multi-hop audio-visual reasoning remains challenging for Omni-LLMs, as relevant evidence is often sparse, temporally dispersed, and distributed across both audio and visual streams. Existing benchmarks provide limited in…

Visual Reasoning

Active Perception Agent for Omnimodal Audio-Video Understanding

2025-12-29 · Keda Tao, Wenjie Du, Bohan Yu, Weiqiang Wang 외 arxiv

Omnimodal large language models have made significant strides in unifying audio and visual modalities; however, they often face challenges in fine-grained cross-modal understanding and have difficulty with multimodal ali…

Response Generation

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning

2026-06-24 · Zhiyuan Han, Beier Zhu, Wenwen Tong, Pengyang Shao 외 arxiv

We find that current emotion-oriented Omni-MLLMs still lack reliable omni-modal perception: they (i) underutilize multimodal cues in their reasoning trajectories and (ii) exhibit unfaithful behavior, often hallucinating …

Reinforcement Learning

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning

2026-01-14 · Dongjie Cheng, Yongqi Li, Zhixin Ma, Hongru Cai 외 arxiv

Multimodal Large Language Models (MLLMs) are making significant progress in multimodal reasoning. Early approaches focus on pure text-based reasoning. More recent studies have incorporated multimodal information into the…

Multimodal ReasoningImage Generation