paper-with-me

홈 › Papers

OMHBench: Benchmarking Balanced and Grounded Omni-Modal Multi-Hop Reasoning

2025-08-22 · Seunghee Kim, Ingyu Bang, Seokgyu Jang, Changhyeon Kim, Sanghwan Bae, Jihun Choi, Richeng Xuan, Taeuk Kim arxiv

Multimodal Large Language Models (MLLMs) have increasingly supported omni-modal processing across text, vision, and speech. However, existing evaluation frameworks for such models suffer from critical limitations, including modality shortcuts and biased reasoning paths. To address these challenges, we propose OMHBench, a novel benchmark designed to rigorously evaluate omni-modal multi-hop reasoning. It consists of 6,144 questions with balanced reasoning paths that are jointly grounded across all three modalities. Extensive evaluation of 13 state-of-the-art models reveals that (1) a large performance gap exists between proprietary and open-source MLLMs and (2) even proprietary models exhibit high sensitivity to reasoning path variations, resulting in asymmetric omni-modal grounding. Notably, models struggle when processing the speech modality, underscoring the need for balanced, multi-hop evaluation of omni-modal intelligence.

📄 PDF Abstract BibTeX arXiv:2508.16198

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants

2026-05-26 · Xudong Lu, Xueying Li, Annan Wang, Yang Bo 외 arxiv

We introduce OmniInteract, a streaming benchmark for real-time omnimodal large language models evaluated through native online inference over audio-visual streams. Unlike offline video understanding or text-prompted stre…

Mathematical Reasoning

OmniACBench: A Benchmark for Evaluating Context-Grounded Acoustic Control in Omni-Modal Models

2026-03-25 · Seunghee Kim, Bumkyu Park, Kyudan Jung, Joosung Lee 외 arxiv

Most testbeds for omni-modal models assess multimodal understanding via textual outputs, leaving it unclear whether these models can properly speak their answers. To study this, we introduce OmniACBench, a benchmark for …

OmniVL-Guard: Towards Unified Vision-Language Forgery Detection and Grounding via Balanced RL

2026-02-11 · Jinjie Shen, Jing Wu, Yaxiong Wang, Lechao Cheng 외 arxiv

Existing forgery detection methods are often limited to uni-modal or bi-modal settings, failing to handle the interleaved text, images, and videos prevalent in real-world misinformation. To bridge this gap, this paper ta…

Reinforcement Learning

OpenOmni: A Collaborative Open Source Tool for Building Future-Ready Multimodal Conversational Agents

2024-08-06 · Qiang Sun, Yuanyi Luo, Sirui Li, Wenxiao Zhang 외

Multimodal conversational agents are highly desirable because they offer natural and human-like interaction. However, there is a lack of comprehensive end-to-end solutions to support collaborative development and benchma…

BenchmarkingRetrieval-augmented GenerationSpeech-to-Text

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos

2026-05-08 · Hengyi Feng, Hao Liang, Mingrui Chen, Bohan Zeng 외 arxiv

Real-world audio-visual understanding requires chaining evidence that is sparse, temporally dispersed, and split across the visual and auditory streams, whereas existing benchmarks largely fail to evaluate this capabilit…

Multimodal Reasoning