paper-with-me

Papers

Evaluating Time Awareness and Cross-modal Active Perception of Large Models via 4D Escape Room Task

2026-03-16 · Yurui Dong, Ziyue Wang, Shuyun Lu, Dairu Liu, Xuechen Liu, Fuwen Luo, Peng Li, Yang Liu arxiv

Multimodal Large Language Models (MLLMs) have recently made rapid progress toward unified Omni models that integrate vision, language, and audio. However, existing environments largely focus on 2D or 3D visual context and vision-language tasks, offering limited support for temporally dependent auditory signals and selective cross-modal integration, where different modalities may provide complementary or interfering information, which are essential capabilities for realistic multimodal reasoning. As a result, whether models can actively coordinate modalities and reason under time-varying, irreversible conditions remains underexplored. To this end, we introduce \textbf{EscapeCraft-4D}, a customizable 4D environment for assessing selective cross-modal perception and time awareness in Omni models. It incorporates trigger-based auditory sources, temporally transient evidence, and location-dependent cues, requiring agents to perform spatio-temporal reasoning and proactive multimodal integration under time constraints. Building on this environment, we curate a benchmark to evaluate corresponding abilities across powerful models. Evaluation results suggest that models struggle with modality bias, and reveal significant gaps in current model's ability to integrate multiple modalities under time constraints. Further in-depth analysis uncovers how multiple modalities interact and jointly influence model decisions in complex multimodal reasoning environments.

📄 PDF Abstract BibTeX arXiv:2603.15467

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Reasoning

Similar Papers 제목 키워드 기반

Evaluating Proactive Risk Awareness of Large Language Models

2026-02-24 · Xuan Luo, Yubin Chen, Zhiyu Hou, Linpu Yu 외 arxiv

As large language models (LLMs) are increasingly embedded in everyday decision-making, their safety responsibilities extend beyond reacting to explicit harmful intent toward anticipating unintended but consequential risk…

MINED: Probing and Updating with Multimodal Time-Sensitive Knowledge for Large Multimodal Models

2025-10-22 · Kailin Jiang, Ning Jiang, Yuntao Du, Yuchen Ren 외 arxiv

Large Multimodal Models (LMMs) encode rich factual knowledge via cross-modal pre-training, yet their static representations struggle to maintain an accurate understanding of time-sensitive factual knowledge. Existing ben…

knowledge editing

PLAICraft: Large-Scale Time-Aligned Vision-Speech-Action Dataset for Embodied AI

2025-05-19 · Yingchen He, Christian D. Weilbach, Martyna E. Wojciechowska, Yuxuan Zhang 외

Advances in deep generative modelling have made it increasingly plausible to train human-level embodied agents. Yet progress has been limited by the absence of large-scale, real-time, multi-modal, and socially interactiv…

BenchmarkingMinecraftObject Recognition

Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs

2025-02-16 · Wenxuan Wang, Xiaoyuan Liu, Kuiyi Gao, Jen-tse Huang 외

Multimodal Large Language Models (MLLMs) have expanded the capabilities of traditional language models by enabling interaction through both text and images. However, ensuring the safety of these models remains a signific…

Benchmarking

IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks

2025-06-19 · Xiaoya Lu, Zeren Chen, Xuhao Hu, Yijin Zhou 외

Flawed planning from VLM-driven embodied agents poses significant safety hazards, hindering their deployment in real-world household tasks. However, existing static, non-interactive evaluation paradigms fail to adequatel…