paper-with-me

홈 › Papers

Seeing Speech and Sound: Distinguishing and Locating Audio Sources in Visual Scenes

2025-01-01 · CVPR 2025 1 · Hyeonggon Ryu, Seongyu Kim, Joon Son Chung, Arda Senocak

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited to handling either speech or non-speech sounds independently, or at best, together but sequentially without mixing. This limitation prevents them from capturing the complexity of real-world audio sources that are often mixed. Our approach introduces a "mix-and-separate" framework with audio-visual alignment objectives that jointly learn correspondence and disentanglement using mixed audio. Through these objectives, our model learns to produce distinct embeddings for each audio type, enabling effective disentanglement and grounding across mixed audio sources.Additionally, we created a new dataset to evaluate simultaneous grounding of mixed audio sources, demonstrating that our model outperforms prior methods. Our approach also achieves state-of-the-art performance in standard segmentation and cross-modal retrieval tasks, highlighting the benefits of our mix-and-separate approach.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalDisentanglementVisual Grounding

Similar Papers 제목 키워드 기반

Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes

2025-03-24 · Hyeonggon Ryu, Seongyu Kim, Joon Son Chung, Arda Senocak

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches ar…

Cross-Modal RetrievalDisentanglementVisual Grounding

Seeing Through Noise: Visually Driven Speaker Separation and Enhancement

2017-08-22 · Aviv Gabbay, Ariel Ephrat, Tavi Halperin, Shmuel Peleg

Isolating the voice of a specific person while filtering out other voices or background noises is challenging when video is shot in noisy environments. We propose audio-visual methods to isolate the voice of a single spe…

Speaker Separation

SEE-2-SOUND: Zero-Shot Spatial Environment-to-Spatial Sound

2024-06-06 · Rishit Dagli, Shivesh Prakash, Robert Wu, Houman Khosravani

Generating combined visual and auditory sensory experiences is critical for the consumption of immersive content. Recent advances in neural generative models have enabled the creation of high-resolution content across mu…

Audio Generation

SeeingSounds: Learning Audio-to-Visual Alignment via Text

2025-10-10 · Simone Carnemolla, Matteo Pennisi, Chiara Russo, Simone Palazzo 외 arxiv

We introduce SeeingSounds, a lightweight and modular framework for audio-to-image generation that leverages the interplay between audio, language, and vision-without requiring any paired audio-visual data or training on …

Image Generation

Seeing Voices: Generating A-Roll Video from Audio with Mirage

2025-06-09 · Aditi Sundararaman, Amogh Adishesha, Andrew Jaegle, Dan Bigioi 외

From professional filmmaking to user-generated content, creators and consumers have long recognized that the power of video depends on the harmonious integration of what we hear (the video's audio track) with what we see…

Speech Synthesistext-to-speechText to SpeechVideo Generation