paper-with-me

Papers

Spatial Audio Motion Understanding and Reasoning

2025-09-18 · Arvind Krishna Sridhar, Yinyi Guo, Erik Visser arxiv

Spatial audio reasoning enables machines to interpret auditory scenes by understanding events and their spatial attributes. In this work, we focus on spatial audio understanding with an emphasis on reasoning about moving sources. First, we introduce a spatial audio encoder that processes spatial audio to detect multiple overlapping events and estimate their spatial attributes, Direction of Arrival (DoA) and source distance, at the frame level. To generalize to unseen events, we incorporate an audio grounding model that aligns audio features with semantic audio class text embeddings via a cross-attention mechanism. Second, to answer complex queries about dynamic audio scenes involving moving sources, we condition a large language model (LLM) on structured spatial attributes extracted by our model. Finally, we introduce a spatial audio motion understanding and reasoning benchmark dataset and demonstrate our framework's performance against the baseline model.

📄 PDF Abstract BibTeX arXiv:2509.14666

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Spatial Audio Question Answering and Reasoning on Dynamic Source Movements

2026-02-18 · Arvind Krishna Sridhar, Yinyi Guo, Erik Visser arxiv

Spatial audio understanding aims to enable machines to interpret complex auditory scenes, particularly when sound sources move over time. In this work, we study Spatial Audio Question Answering (Spatial AQA) with a focus…

Question Answering

AudioMotionBench: Evaluating Auditory Motion Perception in Audio LLMs

2025-11-17 · Zhe Sun, Yujun Cai, Jiayu Yao, Yiwei Wang arxiv

Large Audio-Language Models (LALMs) have recently shown impressive progress in speech recognition, audio captioning, and auditory question answering. Yet, whether these models can perceive spatial dynamics, particularly …

Speech RecognitionQuestion AnsweringSpatial ReasoningAudio captioning

Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding

2026-06-09 · Zhiyuan Zhu, Yixuan Chen, Yiwen Shao, Wenxiang Guo 외 arxiv

Recent multimodal large language models mainly process audio as monaural signals, thereby discarding the spatial cues contained in spatial audio for sound localization, spatial relation reasoning, and spatial scene under…

Scene UnderstandingQuestion AnsweringSpatial Reasoning

Query-Guided Spatial-Temporal-Frequency Interaction for Music Audio-Visual Question Answering

2026-01-27 · Kun Li, Michael Ying Yang, Sami Sebastian Brandt arxiv

Audio--Visual Question Answering (AVQA) is a challenging multimodal task that requires jointly reasoning over audio, visual, and textual information in a given video to answer natural language questions. Inspired by rece…

Audio-visual Question Answering

SPUR: A Plug-and-Play Framework for Integrating Spatial Audio Understanding and Reasoning into Large Audio-Language Models

2025-11-10 · S Sakshi, Vaibhavi Lokegaonkar, Neil Zhang, Ramani Duraiswami 외 arxiv

Spatial perception is central to auditory intelligence, enabling accurate understanding of real-world acoustic scenes and advancing human-level perception of the world around us. While recent large audio-language models …

Spatial Reasoning