paper-with-me

홈 › Papers

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

2025-02-06 · Jack Hong, Shilin Yan, Jiayin Cai, XiaoLong Jiang, Yao Hu, Weidi Xie

In this paper, we introduce WorldSense, the first benchmark to assess the multi-modal video understanding, that simultaneously encompasses visual, audio, and text inputs. In contrast to existing benchmarks, our WorldSense has several features: (i) collaboration of omni-modality, we design the evaluation tasks to feature a strong coupling of audio and video, requiring models to effectively utilize the synergistic perception of omni-modality; (ii) diversity of videos and tasks, WorldSense encompasses a diverse collection of 1,662 audio-visual synchronised videos, systematically categorized into 8 primary domains and 67 fine-grained subcategories to cover the broad scenarios, and 3,172 multi-choice QA pairs across 26 distinct tasks to enable the comprehensive evaluation; (iii) high-quality annotations, all the QA pairs are manually labeled by 80 expert annotators with multiple rounds of correction to ensure quality. Based on our WorldSense, we extensively evaluate various state-of-the-art models. The experimental results indicate that existing models face significant challenges in understanding real-world scenarios (48.0% best accuracy). We hope our WorldSense can provide a platform for evaluating the ability in constructing and understanding coherent contexts from omni-modality.

📄 PDF Abstract BibTeX arXiv:2502.04326

Code (0)

등록된 구현이 없습니다.

Tasks

Video Understanding

Similar Papers 제목 키워드 기반

OmniRAG-Agent: Agentic Omnimodal Reasoning for Low-Resource Long Audio-Video Question Answering

2026-02-03 · Yifan Zhu, Xinyu Mu, Tao Feng, Zhonghong Ou 외 arxiv

Long-horizon omnimodal question answering answers questions by reasoning over text, images, audio, and video. Despite recent progress on OmniLLMs, low-resource long audio-video QA still suffers from costly dense encoding…

Video Question Answering

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models

2026-05-12 · Yuchen Deng, Zidang Cai, Hai-Tao Zheng, Jie Wang 외 arxiv

Omnimodal large language models (Omni-LLMs) show strong capability in audio-video understanding, but their practical deployment remains limited by high inference cost of long video streams and dense audio sequences. Desp…

Demographic and Linguistic Bias Evaluation in Omnimodal Language Models

2026-04-11 · Alaa Elobaid arxiv

This paper provides a comprehensive evaluation of demographic and linguistic biases in omnimodal language models that process text, images, audio, and video within a single framework. Although these models are being wide…

Language IdentificationActivity Recognition

DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression

2026-03-15 · Bingzhou Li, Tao Huang arxiv

Omnimodal large language models (OmniLLMs) jointly process audio and visual streams, but the resulting long multimodal token sequences make inference prohibitively expensive. Existing compression methods typically rely o…

OpenOmni: Large Language Models Pivot Zero-shot Omnimodal Alignment across Language with Real-time Self-Aware Emotional Speech Synthesis

2025-01-08 · Run Luo, Ting-En Lin, Haonan Zhang, Yuchuan Wu 외

Recent advancements in omnimodal learning have been achieved in understanding and generation across images, text, and speech, though mainly within proprietary models. Limited omnimodal datasets and the inherent challenge…

DecoderEmotional Speech SynthesisLanguage ModelingLanguage Modelling+2