paper-with-me

홈 › Papers

Omni-Modal Dissonance Benchmark: Systematically Breaking Modality Consensus to Probe Robustness and Calibrated Abstention

2026-03-28 · Zabir Al Nazi, Shubhashis Roy Dipta, Md Rizwan Parvez arxiv

Existing omni-modal benchmarks attempt to measure modality-specific contributions, but their measurements are confounded: naturally co-occurring modalities carry correlated yet unequal information, making it unclear whether results reflect true modality reliance or information asymmetry. We introduce OMD-Bench, where all modalities are initially congruent - each presenting the same anchor, an object or event independently perceivable through video, audio, and text - which we then systematically corrupt to isolate each modality's contribution. We also evaluate calibrated abstention: whether models appropriately refrain from answering when evidence is conflicting. The benchmark comprises 4,080 instances spanning 27 anchors across eight corruption conditions. Evaluating ten omni-modal models under zero-shot and chain-of-thought prompting, we find that models over-abstain when two modalities are corrupted yet under-abstain severely when all three are, while maintaining high confidence (~60-100%) even under full corruption. Chain-of-thought prompting improves abstention alignment with human judgment but amplifies overconfidence rather than mitigating it. OMD-Bench provides a diagnostic benchmark for diagnosing modality reliance, robustness to cross-modal inconsistency, and uncertainty calibration in omni-modal systems.

📄 PDF Abstract BibTeX arXiv:2603.27187

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniPlay: Benchmarking Omni-Modal Models on Omni-Modal Game Playing

2025-08-06 · Fuqing Bie, Shiyu Huang, Xijia Tao, Zhiqin Fang 외 arxiv

While generalist foundation models like Gemini and GPT-4o demonstrate impressive multi-modal competence, existing evaluations fail to test their intelligence in dynamic, interactive worlds. Static benchmarks lack agency,…

OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora

2026-08-26 · Hao Chen, Yumin Lin, Nadila Yushanjiang, Xin Lin 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated strong abilities in solving diverse visual and textual reasoning tasks. However, their development in the physics domain is significantly hindered by the lack of…

OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments

2026-04-03 · Felix Henry, Xiaochen Lin, Jiangyou Zhu, Yangfan 외 arxiv

Current benchmarks for graphical user interface (GUI) agents predominantly rely on static screenshots. However, real-world smartphone interaction routinely requires agents to process transient audio cues and temporal vid…

Empirical Evaluation of Topic Zero- and Few-Shot Learning for Stance Dissonance Detection

2021-11-16 · ACL ARR November 2021 11 · Anonymous

We address stance dissonance detection, the task of detecting conflicting stance between two input statements. Computational models for traditional stance detection have typically been trained to indicate pro/cons for a …

Few-Shot LearningStance Detection

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

2025-02-06 · Jack Hong, Shilin Yan, Jiayin Cai, XiaoLong Jiang 외

In this paper, we introduce WorldSense, the first benchmark to assess the multi-modal video understanding, that simultaneously encompasses visual, audio, and text inputs. In contrast to existing benchmarks, our WorldSens…

Video Understanding