paper-with-me

홈 › Papers

Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes

2025-03-24 · Hyeonggon Ryu, Seongyu Kim, Joon Son Chung, Arda Senocak

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited to handling either speech or non-speech sounds independently, or at best, together but sequentially without mixing. This limitation prevents them from capturing the complexity of real-world audio sources that are often mixed. Our approach introduces a 'mix-and-separate' framework with audio-visual alignment objectives that jointly learn correspondence and disentanglement using mixed audio. Through these objectives, our model learns to produce distinct embeddings for each audio type, enabling effective disentanglement and grounding across mixed audio sources. Additionally, we created a new dataset to evaluate simultaneous grounding of mixed audio sources, demonstrating that our model outperforms prior methods. Our approach also achieves comparable or better performance in standard segmentation and cross-modal retrieval tasks, highlighting the benefits of our mix-and-separate approach.

📄 PDF Abstract BibTeX arXiv:2503.18880

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalDisentanglementVisual Grounding

Similar Papers 제목 키워드 기반

Seeing Speech and Sound: Distinguishing and Locating Audio Sources in Visual Scenes

2025-01-01 · CVPR 2025 1 · Hyeonggon Ryu, Seongyu Kim, Joon Son Chung, Arda Senocak

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approache…

Cross-Modal RetrievalDisentanglementVisual Grounding

Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds

2020-11-02 · ICLR 2021 1 · Efthymios Tzinis, Scott Wisdom, Aren Jansen, Shawn Hershey 외

Recent progress in deep learning has enabled many advances in sound separation and visual scene understanding. However, extracting sound sources which are apparent in natural videos remains an open problem. In this work,…

Scene Understanding

AudioSR: Versatile Audio Super-resolution at Scale

2023-09-13 · Haohe Liu, Ke Chen, Qiao Tian, Wenwu Wang 외

Audio super-resolution is a fundamental task that predicts high-frequency components for low-resolution audio, enhancing audio quality in digital applications. Previous methods have limitations such as the limited scope …

Audio Super-ResolutionSuper-Resolution

Evaluating robustness of You Only Hear Once(YOHO) Algorithm on noisy audios in the VOICe Dataset

2021-11-01 · Soham Tiwari, Kshitiz Lakhotia, Manjunath Mulimani

Sound event detection (SED) in machine listening entails identifying the different sounds in an audio file and identifying the start and end time of a particular sound event in the audio. SED finds use in various applica…

Event DetectionRetrievalSound Event Detectionspeech-recognition+1

Separate Anything You Describe

2023-08-09 · Xubo Liu, Qiuqiang Kong, Yan Zhao, Haohe Liu 외

Language-queried audio source separation (LASS) is a new paradigm for computational auditory scene analysis (CASA). LASS aims to separate a target sound from an audio mixture given a natural language query, which provide…

Audio Source SeparationNatural Language QueriesSpeech EnhancementZero-shot Generalization