paper-with-me

홈 › Papers

I see what you hear: a vision-inspired method to localize words

2022-10-24 · Mohammad Samragh, Arnav Kundu, Ting-yao Hu, Minsik Cho, Aman Chadha, Ashish Shrivastava, Oncel Tuzel, Devang Naik

This paper explores the possibility of using visual object detection techniques for word localization in speech data. Object detection has been thoroughly studied in the contemporary literature for visual data. Noting that an audio can be interpreted as a 1-dimensional image, object localization techniques can be fundamentally useful for word localization. Building upon this idea, we propose a lightweight solution for word detection and localization. We use bounding box regression for word localization, which enables our model to detect the occurrence, offset, and duration of keywords in a given audio stream. We experiment with LibriSpeech and train a model to localize 1000 words. Compared to existing work, our method reduces model size by 94%, and improves the F1 score by 6.5\%.

📄 PDF Abstract BibTeX arXiv:2210.13567

Code (0)

등록된 구현이 없습니다.

Tasks

Objectobject-detectionObject DetectionObject Localization

Similar Papers 제목 키워드 기반

What Do Deepfake Speech Detectors Actually Hear?

2026-06-09 · Vojtěch Staněk, Veronika Jirmusová, Anton Firc, Kamil Malinka 외 arxiv

Deepfake speech detectors often output a single score without explaining why an audio sample is flagged, where in the signal the evidence lies, or what cues drive the decision. We propose an audio-native explainability p…

Learning to Generate Grounded Visual Captions without Localization Supervision

2019-06-01 · Chih-Yao Ma, Yannis Kalantidis, Ghassan AlRegib, Peter Vajda 외

When automatically generating a sentence description for an image or video, it often remains unclear how well the generated caption is grounded, that is whether the model uses the correct image regions to output particul…

Image CaptioningLanguage ModellingSentenceVideo Captioning

Learning to Generate Grounded Visual Captions without Localization Supervision

2020-08-01 · ECCV 2020 8 · Chih-Yao Ma, Yannis Kalantidis, Ghassan AlRegib, Peter Vajda 외

When automatically generating a sentence description for an image or video, it often remains unclear how well the generated caption is grounded, that is whether the model uses the correct image regions to output particul…

Image CaptioningLanguage ModellingSentenceVideo Captioning

Visual Words for Automatic Lip-Reading

2014-09-17 · Ahmad Basheer Hassanat

Lip reading is used to understand or interpret speech without hearing it, a technique especially mastered by people with hearing difficulties. The ability to lip read enables a person with a hearing impairment to communi…

Lip Readingspeech-recognitionSpeech RecognitionVisual Speech Recognition

Audio Geolocation: A Natural Sounds Benchmark

2025-05-24 · Mustafa Chasmai, Wuao Liu, Subhransu Maji, Grant van Horn

Can we determine someone's geographic location purely from the sounds they hear? Are acoustic signals enough to localize within a country, state, or even city? We tackle the challenge of global-scale audio geolocation, f…