paper-with-me

홈 › Papers

You Only Speak Once to See

2024-09-27 · Wenhao Yang, Jianguo Wei, Wenhuan Lu, Lei LI

Grounding objects in images using visual cues is a well-established approach in computer vision, yet the potential of audio as a modality for object recognition and grounding remains underexplored. We introduce YOSS, "You Only Speak Once to See," to leverage audio for grounding objects in visual scenes, termed Audio Grounding. By integrating pre-trained audio models with visual models using contrastive learning and multi-modal alignment, our approach captures speech commands or descriptions and maps them directly to corresponding objects within images. Experimental results indicate that audio guidance can be effectively applied to object grounding, suggesting that incorporating audio guidance may enhance the precision and robustness of current object grounding methods and improve the performance of robotic systems and computer vision applications. This finding opens new possibilities for advanced object recognition, scene understanding, and the development of more intuitive and capable robotic systems.

📄 PDF Abstract BibTeX arXiv:2409.18372

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningObjectObject RecognitionScene Understanding

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

End-to-end losses based on speaker basis vectors and all-speaker hard negative mining for speaker verification

2019-02-07 · Hee-Soo Heo, Jee-weon Jung, IL-Ho Yang, Sung-Hyun Yoon 외

In recent years, speaker verification has primarily performed using deep neural networks that are trained to output embeddings from input features such as spectrograms or Mel-filterbank energies. Studies that design vari…

AllMetric LearningSpeaker Verification

$\mathsf{OPA}$: One-shot Private Aggregation with Single Client Interaction and its Applications to Federated Learning

2024-10-29 · Harish Karthikeyan, Antigoni Polychroniadou

Our work aims to minimize interaction in secure computation due to the high cost and challenges associated with communication rounds, particularly in scenarios with many clients. In this work, we revisit the problem of s…

Federated LearningPrivacy Preserving

Weakly Supervised Training of Speaker Identification Models

2018-06-22 · Martin Karu, Tanel Alumäe

We propose an approach for training speaker identification models in a weakly supervised manner. We concentrate on the setting where the training data consists of a set of audio recordings and the speaker annotation is p…

speaker-diarizationSpeaker DiarizationSpeaker Identification

On Feature Importance and Interpretability of Speaker Representations

2023-10-19 · Frederik Rautenberg, Michael Kuhlmann, Jana Wiechmann, Fritz Seebauer 외

Unsupervised speech disentanglement aims at separating fast varying from slowly varying components of a speech signal. In this contribution, we take a closer look at the embedding vector representing the slowly varying s…

DisentanglementFeature Importance

Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation

2018-04-10 · Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel 외

We present a joint audio-visual model for isolating a single speech signal from a mixture of sounds such as other speakers and background noise. Solving this task using only audio as input is extremely challenging and do…

Speech Separation