paper-with-me

홈 › Papers

iQuery: Instruments as Queries for Audio-Visual Sound Separation

2022-12-07 · CVPR 2023 1 · Jiaben Chen, Renrui Zhang, Dongze Lian, Jiaqi Yang, Ziyao Zeng, Jianbo Shi

Current audio-visual separation methods share a standard architecture design where an audio encoder-decoder network is fused with visual encoding features at the encoder bottleneck. This design confounds the learning of multi-modal feature encoding with robust sound decoding for audio separation. To generalize to a new instrument: one must finetune the entire visual and audio network for all musical instruments. We re-formulate visual-sound separation task and propose Instrument as Query (iQuery) with a flexible query expansion mechanism. Our approach ensures cross-modal consistency and cross-instrument disentanglement. We utilize "visually named" queries to initiate the learning of audio queries and use cross-modal attention to remove potential sound source interference at the estimated waveforms. To generalize to a new instrument or event class, drawing inspiration from the text-prompt design, we insert an additional query as an audio prompt while freezing the attention mechanism. Experimental results on three benchmarks demonstrate that our iQuery improves audio-visual sound source separation performance.

📄 PDF Abstract BibTeX arXiv:2212.03814

Code (1)

jiabenchen/iquery 공식 구현 pytorch

Tasks

DecoderDisentanglement

Similar Papers 제목 키워드 기반

Discovering Sounding Objects by Audio Queries for Audio Visual Segmentation

2023-09-18 · Shaofei Huang, Han Li, Yuqing Wang, Hongji Zhu 외

Audio visual segmentation (AVS) aims to segment the sounding objects for each frame of a given video. To distinguish the sounding objects from silent ones, both audio-visual semantic correspondence and temporal interacti…

ObjectSemantic correspondence

DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information

2024-09-18 · Shota Nakada, Taichi Nishimura, Hokuto Munakata, Masayoshi Kondo 외

Current audio-visual representation learning can capture rough object categories (e.g., ``animals'' and ``instruments''), but it lacks the ability to recognize fine-grained details, such as specific categories like ``dog…

ObjectRepresentation LearningRetrieval

Separate Anything You Describe

2023-08-09 · Xubo Liu, Qiuqiang Kong, Yan Zhao, Haohe Liu 외

Language-queried audio source separation (LASS) is a new paradigm for computational auditory scene analysis (CASA). LASS aims to separate a target sound from an audio mixture given a natural language query, which provide…

Audio Source SeparationNatural Language QueriesSpeech EnhancementZero-shot Generalization

OmniQuery: Contextually Augmenting Captured Multimodal Memory to Enable Personal Question Answering

2024-09-12 · Jiahao Nick Li, Zhuohao Jerry Zhang, Jiaju Ma

People often capture memories through photos, screenshots, and videos. While existing AI-based tools enable querying this data using natural language, they only support retrieving individual pieces of information like ce…

Language ModelingLanguage ModellingLarge Language ModelQuestion Answering+1

MoXaRt: Audio-Visual Object-Guided Sound Interaction for XR

2026-03-11 · Tianyu Xu, Sieun Kim, Qianhui Zheng, Ruoyu Xu 외 arxiv

In Extended Reality (XR), complex acoustic environments often overwhelm users, compromising both scene awareness and social engagement due to entangled sound sources. We introduce MoXaRt, a real-time XR system that uses …