paper-with-me

홈 › Papers

The World is Not Mono: Enabling Spatial Understanding in Large Audio-Language Models

2026-01-06 · Yuhuan You, Lai Wei, Xihong Wu, Tianshu Qu arxiv

Large audio-language models have made rapid progress in recognizing what is present in an audio clip, but spatial audio-language understanding still lacks a clear task interface. A model must also decide where sound events occur, which semantic and spatial attributes belong to the same auditory object, how multiple objects are arranged, and whether a scene-level answer is physically plausible. We formalize this capability as audio scene analysis (ASA), a three-level problem spanning atomic perception, relational integration, and cognitive reasoning. We propose The World is Not Mono (TWNM), a framework that equips audio-language models with explicit spatial evidence. TWNM uses physically grounded First-Order Ambisonics (FOA) simulation for controllable supervision, learns slot-regularized spatial representations from multichannel audio, fuses them with semantic audio features, and trains with a progressive curriculum ending in preference optimization over metadata-derived answers and auxiliary format/evidence rewards. To operationalize ASA, we build a controlled benchmark from scene metadata, covering localization, attribute binding, spatial comparison, scene abduction, and counterfactual reasoning. On this benchmark, TWNM achieves 70.8% overall accuracy, 66.4% on spatial-family tasks, and 79.76% on mixed L3 scene-level multiple-choice QA. We also audit monaural and binaural reference systems as diagnostic references with explicit audit labels, since they differ in spatial input, training interface, and output format. The supported claim is that a clearly defined ASA hierarchy, FOA-conditioned spatial representations, and metadata-grounded training enable controlled, auditable spatial audio-language reasoning, with STARSS23 providing a limited real-recording diagnostic.

📄 PDF Abstract BibTeX arXiv:2601.02954

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OpenMonoGS-SLAM: Monocular Gaussian Splatting SLAM with Open-set Semantics

2025-12-09 · Jisang Yoo, Gyeongjin Kang, Hyun-kyu Ko, Hyeonwoo Yu 외 arxiv

Simultaneous Localization and Mapping (SLAM) is a foundational component in robotics, AR/VR, and autonomous systems. With the rising focus on spatial AI in recent years, combining SLAM with semantic understanding has bec…

Self-Supervised Learning

Beyond Human Perception: Understanding Multi-Object World from Monocular View

2025-01-01 · CVPR 2025 1 · Keyu Guo, Yongle Huang, ShiJie Sun, XiangYu Song 외

Language and binocular vision play a crucial role in human understanding of the world. Advancements in artificial intelligence have also made it possible for machines to develop 3D perception capabilities essential f…

3D visual groundingDenoisingScene UnderstandingVisual Grounding

Seeing World Dynamics in a Nutshell

2025-02-05 · Qiuhong Shen, Xuanyu Yi, Mingbao Lin, Hanwang Zhang 외

We consider the problem of efficiently representing casually captured monocular videos in a spatially- and temporally-coherent manner. While existing approaches predominantly rely on 2D/2.5D techniques treating videos as…

Video Reconstruction

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images

2025-11-24 · Qirui Wang, Jingyi He, Yining Pan, Si Yong Yeo 외 arxiv

Spatial reasoning (SR), the ability to infer 3D spatial information from 2D inputs, is essential for real-world applications such as embodied AI and autonomous driving. However, existing research primarily focuses on ind…

Autonomous DrivingSpatial Reasoning

M2H-MX: Multi-Task Semantic and Geometric Perception for Real-Time Monocular 3D Scene Graph Construction

2026-03-31 · U. V. B. L. Udugama, George Vosselman, Francesco Nex arxiv

Monocular cameras are attractive for robotic perception due to their low cost and ease of deployment, yet achieving reliable real-time spatial understanding from a single image stream remains challenging. While recent mu…