paper-with-me

Papers

Looking Enhances Listening: Recovering Missing Speech Using Images

2020-02-13 · Tejas Srinivasan, Ramon Sanabria, Florian Metze

Speech is understood better by using visual context; for this reason, there have been many attempts to use images to adapt automatic speech recognition (ASR) systems. Current work, however, has shown that visually adapted ASR models only use images as a regularization signal, while completely ignoring their semantic content. In this paper, we present a set of experiments where we show the utility of the visual modality under noisy conditions. Our results show that multimodal ASR models can recover words which are masked in the input acoustic signal, by grounding its transcriptions using the visual representations. We observe that integrating visual context can result in up to 35% relative improvement in masked word recovery. These results demonstrate that end-to-end multimodal ASR systems can become more robust to noise by leveraging the visual context.

📄 PDF Abstract BibTeX arXiv:2002.05639

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

What Did I Just Say? Self-Listening for Full-Duplex Speech Models

2026-09-04 · Xuanning Zhou, Junyi Ao, Xiaotong Liu, Tom Ko 외 hf

Full-duplex spoken language models can listen and speak simultaneously, enabling them to handle interruptions and backchannels in human conversation. However, text generation, speech synthesis, and audio playback proceed…

Speech SynthesisText Generation

Joint Noise Reduction and Listening Enhancement for Full-End Speech Enhancement

2022-03-22 · Haoyu Li, Yun Liu, Junichi Yamagishi

Speech enhancement (SE) methods mainly focus on recovering clean speech from noisy input. In real-world speech communication, however, noises often exist in not only speaker but also listener environments. Although SE me…

Speech Enhancement

Looking and Listening Inside and Outside: Multimodal Artificial Intelligence Systems for Driver Safety Assessment and Intelligent Vehicle Decision-Making

2026-02-07 · Ross Greer, Laura Fleig, Maitrayee Keskar, Erika Maquiling 외 arxiv

The looking-in-looking-out (LILO) framework has enabled intelligent vehicle applications that understand both the outside scene and the driver state to improve safety outcomes, with examples in smart airbag deployment, t…

Scene Understanding

TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation

2025-12-23 · Ji-Hoon Kim, Junseok Ahn, Doyeop Kwak, Joon Son Chung 외 arxiv

The objective of this paper is to jointly synthesize interactive videos and conversational speech from text and reference images. With the ultimate goal of building human-like conversational systems, recent studies have …

Dialogue Generation

Neural Speech Tracking in a Virtual Acoustic Environment: Audio-Visual Benefit for Unscripted Continuous Speech

2025-01-14 · Mareike Daeglau, Juergen Otten, Giso Grimm, Bojana Mirkovic 외

The audio visual benefit in speech perception, where congruent visual input enhances auditory processing, is well documented across age groups, particularly in challenging listening conditions and among individuals with …

EEG