paper-with-me

홈 › Papers

MoXaRt: Audio-Visual Object-Guided Sound Interaction for XR

2026-03-11 · Tianyu Xu, Sieun Kim, Qianhui Zheng, Ruoyu Xu, Tejasvi Ravi, Anuva Kulkarni, Katrina Passarella-Ward, Junyi Zhu, Adarsh Kowdle arxiv

In Extended Reality (XR), complex acoustic environments often overwhelm users, compromising both scene awareness and social engagement due to entangled sound sources. We introduce MoXaRt, a real-time XR system that uses audio-visual cues to separate these sources and enable fine-grained sound interaction. MoXaRt's core is a cascaded architecture that performs coarse, audio-only separation in parallel with visual detection of sources (e.g., faces, instruments). These visual anchors then guide refinement networks to isolate individual sources, separating complex mixes of up to 5 concurrent sources (e.g., 2 voices + 3 instruments) with ~2 second processing latency. We validate MoXaRt through a technical evaluation on a new dataset of 30 one-minute recordings featuring concurrent speech and music, and a 22-participant user study. Empirical results indicate that our system significantly enhances speech intelligibility, yielding a 36.2% (p < 0.01) increase in listening comprehension within adversarial acoustic environments while substantially reducing cognitive load (p < 0.001), thereby paving the way for more perceptive and socially adept XR experiences.

📄 PDF Abstract BibTeX arXiv:2603.10465

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Co-Separating Sounds of Visual Objects

2019-04-16 · ICCV 2019 10 · Ruohan Gao, Kristen Grauman

Learning how objects sound from video is challenging, since they often heavily overlap in a single audio channel. Current methods for visually-guided audio source separation sidestep the issue by training with artificial…

Audio DenoisingAudio Source SeparationDenoising

Weakly-supervised Audio-visual Sound Source Detection and Separation

2021-03-25 · Tanzila Rahman, Leonid Sigal

Learning how to localize and separate individual object sounds in the audio channel of the video is a difficult task. Current state-of-the-art methods predict audio masks from artificially mixed spectrograms, known as Mi…

Audio Source SeparationDenoisingObjectSegmentation+3

Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment

2025-03-17 · CVPR 2025 1 · Chen Liu, Peike Li, Liying Yang, Dadong Wang 외

Accurately localizing audible objects based on audio-visual cues is the core objective of audio-visual segmentation. Most previous methods emphasize spatial or temporal multi-modal modeling, yet overlook challenges from …

Contrastive Learning

Audio-Synchronized Visual Animation

2024-03-08 · Lin Zhang, Shentong Mo, Yijing Zhang, Pedro Morgado

Current visual generation methods can produce high quality videos guided by texts. However, effectively controlling object dynamics remains a challenge. This work explores audio as a cue to generate temporally synchroniz…

LAVSS: Location-Guided Audio-Visual Spatial Audio Separation

2023-10-31 · Yuxin Ye, Wenming Yang, Yapeng Tian

Existing machine learning research has achieved promising results in monaural audio-visual separation (MAVS). However, most MAVS methods purely consider what the sound source is, not where it is located. This can be a pr…