paper-with-me

홈 › Papers

SparseVSR: Lightweight and Noise Robust Visual Speech Recognition

2023-07-10 · Adriana Fernandez-Lopez, Honglie Chen, Pingchuan Ma, Alexandros Haliassos, Stavros Petridis, Maja Pantic

Recent advances in deep neural networks have achieved unprecedented success in visual speech recognition. However, there remains substantial disparity between current methods and their deployment in resource-constrained devices. In this work, we explore different magnitude-based pruning techniques to generate a lightweight model that achieves higher performance than its dense model equivalent, especially under the presence of visual noise. Our sparse models achieve state-of-the-art results at 10% sparsity on the LRS3 dataset and outperform the dense equivalent up to 70% sparsity. We evaluate our 50% sparse model on 7 different visual noise types and achieve an overall absolute improvement of more than 2% WER compared to the dense equivalent. Our results confirm that sparse networks are more resistant to noise than dense networks.

📄 PDF Abstract BibTeX arXiv:2307.04552

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech RecognitionVisual Speech Recognition

Methods 이 논문이 사용한 방법론

Pruning 설명 없음

Similar Papers 제목 키워드 기반

Audio Lottery: Speech Recognition Made Ultra-Lightweight, Noise-Robust, and Transferable

2021-09-29 · ICLR 2022 4 · Shaojin Ding, Tianlong Chen, Zhangyang Wang

Lightweight speech recognition models have seen explosive demands owing to a growing amount of speech-interactive features on mobile devices. Since designing such systems from scratch is non-trivial, practitioners typica…

speech-recognitionSpeech Recognition

Visual Context-driven Audio Feature Enhancement for Robust End-to-End Audio-Visual Speech Recognition

2022-07-13 · Joanna Hong, Minsu Kim, Daehun Yoo, Yong Man Ro

This paper focuses on designing a noise-robust end-to-end Audio-Visual Speech Recognition (AVSR) system. To this end, we propose Visual Context-driven Audio Feature Enhancement module (V-CAFE) to enhance the input noisy …

Audio-Visual Speech RecognitionDecoderNoisy Speech Recognitionspeech-recognition+2

Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech Recognition

2023-06-18 · Yuchen Hu, Ruizhe Li, Chen Chen, Chengwei Qin 외

Audio-visual speech recognition (AVSR) provides a promising solution to ameliorate the noise-robustness of audio-only speech recognition with visual information. However, most existing efforts still focus on audio modali…

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

AVFormer: Injecting Vision into Frozen Speech Models for Zero-Shot AV-ASR

2023-03-29 · CVPR 2023 1 · Paul Hongsuck Seo, Arsha Nagrani, Cordelia Schmid

Audiovisual automatic speech recognition (AV-ASR) aims to improve the robustness of a speech recognition system by incorporating visual information. Training fully supervised multimodal models for this task from scratch,…

Automatic Speech RecognitionDomain AdaptationRobust Speech Recognitionspeech-recognition+1

Visual-Aware Speech Recognition for Noisy Scenarios

2025-04-09 · Lakshmipathi Balaji, Karan Singla

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Vi…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+2