paper-with-me

홈 › Papers

ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?

2025-10-13 · Liu Yang, Huiyu Duan, Ran Tao, Juntao Cheng, Sijing Wu, Yunhao Li, Jing Liu, Xiongkuo Min, Guangtao Zhai arxiv

Omnidirectional images (ODIs) provide full 360x180 view which are widely adopted in VR, AR and embodied intelligence applications. While multi-modal large language models (MLLMs) have demonstrated remarkable performance on conventional 2D image and video understanding benchmarks, their ability to comprehend the immersive environments captured by ODIs remains largely unexplored. To address this gap, we first present ODI-Bench, a novel comprehensive benchmark specifically designed for omnidirectional image understanding. ODI-Bench contains 2,000 high-quality omnidirectional images and over 4,000 manually annotated question-answering (QA) pairs across 10 fine-grained tasks, covering both general-level and spatial-level ODI understanding. Extensive experiments are conducted to benchmark 20 representative MLLMs, including proprietary and open-source models, under both close-ended and open-ended settings. Experimental results reveal that current MLLMs still struggle to capture the immersive context provided by ODIs. To this end, we further introduce Omni-CoT, a training-free method which significantly enhances MLLMs' comprehension ability in the omnidirectional environment through chain-of-thought reasoning across both textual information and visual cues. Both the benchmark and the code will be released at https://github.com/ylylyl-sjtu/ODI-Bench.

📄 PDF Abstract BibTeX arXiv:2510.11549

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Omnidirectional Spatial Modeling from Correlated Panoramas

2025-09-02 · Xinshen Zhang, Tongxi Fu, Xu Zheng arxiv

Omnidirectional scene understanding is vital for various downstream applications, such as embodied AI, autonomous driving, and immersive environments, yet remains challenging due to geometric distortion and complex spati…

Visual Question AnsweringScene UnderstandingAutonomous Driving

Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method

2025-05-20 · Xinshen Zhang, Zhen Ye, Xu Zheng

Omnidirectional images (ODIs), with their 360{\deg} field of view, provide unparalleled spatial awareness for immersive applications like augmented reality and embodied AI. However, the capability of existing multi-modal…

HallucinationObject LocalizationQuestion AnsweringVisual Question Answering

Dense360: Dense Understanding from Omnidirectional Panoramas

2025-06-17 · Yikang Zhou, Tao Zhang, Dizhe Zhang, Shunping Ji 외

Multimodal Large Language Models (MLLMs) require comprehensive visual inputs to achieve dense understanding of the physical world. While existing MLLMs demonstrate impressive world understanding capabilities through limi…

ERP

SphericalDreamer: Generating Navigable Immersive 3D Worlds with Panorama Fusion

2026-05-19 · Antoine Schnepf, Karim Kassab, Flavian Vasile, Andrew Comport arxiv

The generation of immersive and navigable 3D environments is increasingly prevalent with the growing adoption of virtual reality and 3D content. However, recent methods face a fundamental limitation: they cannot produce …

Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images

2026-05-12 · Yuangong Chen, Wai Keung Wong, Jiaxing Li, Ioannis Patras 외 arxiv

Multimodal Large Language Models (MLLMs) show strong visual perception, yet remain limited in reasoning about space under changing viewpoints. We study this challenge as Perspective-Conditioned Spatial Reasoning (PCSR) i…

Spatial ReasoningObject Counting