paper-with-me

홈 › Papers

Enhancing CTC-Based Visual Speech Recognition

2024-09-11 · Hendrik Laux, Anke Schmeink

This paper presents LiteVSR2, an enhanced version of our previously introduced efficient approach to Visual Speech Recognition (VSR). Building upon our knowledge distillation framework from a pre-trained Automatic Speech Recognition (ASR) model, we introduce two key improvements: a stabilized video preprocessing technique and feature normalization in the distillation process. These improvements yield substantial performance gains on the LRS2 and LRS3 benchmarks, positioning LiteVSR2 as the current best CTC-based VSR model without increasing the volume of training data or computational resources utilized. Furthermore, we explore the scalability of our approach by examining performance metrics across varying model complexities and training data volumes. LiteVSR2 maintains the efficiency of its predecessor while significantly enhancing accuracy, thereby demonstrating the potential for resource-efficient advancements in VSR technology.

📄 PDF Abstract BibTeX arXiv:2409.07210

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Knowledge Distillationspeech-recognitionSpeech RecognitionVisual Speech Recognition

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

Continuous Speech Recognition using EEG and Video

2019-12-16 · Gautam Krishna, Mason Carnahan, Co Tran, Ahmed H. Tewfik

In this paper we investigate whether electroencephalography (EEG) features can be used to improve the performance of continuous visual speech recognition systems. We implemented a connectionist temporal classification (C…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)EEGElectroencephalogram (EEG)+4

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition

2025-02-11 · Sungnyun Kim, Kangwook Jang, Sangmin Bae, Sungwoo Cho 외

Audio-visual speech recognition (AVSR) has become critical for enhancing speech recognition in noisy environments by integrating both auditory and visual modalities. However, existing AVSR systems struggle to scale up wi…

Audio-Visual Speech RecognitionComputational EfficiencyMixture-of-ExpertsRobust Speech Recognition+3

Enhancing Audiovisual Speech Recognition through Bifocal Preference Optimization

2024-12-26 · Yihan Wu, Yichen Lu, Yifan Peng, Xihua Wang 외

Audiovisual Automatic Speech Recognition (AV-ASR) aims to improve speech recognition accuracy by leveraging visual signals. It is particularly challenging in unconstrained real-world scenarios across various domains due …

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition

2026-01-18 · Linzhi Wu, Xingyu Zhang, Hao Yuan, Yakun Zhang 외 arxiv

Audio-visual speech recognition (AVSR) typically improves recognition accuracy in noisy environments by integrating noise-immune visual cues with audio signals. Nevertheless, high-noise audio inputs are prone to introduc…

Audio-Visual Speech RecognitionSpeech Enhancement

VILAS: Exploring the Effects of Vision and Language Context in Automatic Speech Recognition

2023-05-31 · Ziyi Ni, Minglun Han, Feilong Chen, Linghui Meng 외

Enhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these works have primarily focused on utilizing v…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition