paper-with-me

홈 › Papers

Lip-Listening: Mixing Senses to Understand Lips using Cross Modality Knowledge Distillation for Word-Based Models

2022-06-05 · Hadeel Mabrouk, Omar Abugabal, Nourhan Sakr, Hesham M. Eraqi

In this work, we propose a technique to transfer speech recognition capabilities from audio speech recognition systems to visual speech recognizers, where our goal is to utilize audio data during lipreading model training. Impressive progress in the domain of speech recognition has been exhibited by audio and audio-visual systems. Nevertheless, there is still much to be explored with regards to visual speech recognition systems due to the visual ambiguity of some phonemes. To this end, the development of visual speech recognition models is crucial given the instability of audio models. The main contributions of this work are i) building on recent state-of-the-art word-based lipreading models by integrating sequence-level and frame-level Knowledge Distillation (KD) to their systems; ii) leveraging audio data during training visual models, a feat which has not been utilized in prior word-based work; iii) proposing the Gaussian-shaped averaging in frame-level KD, as an efficient technique that aids the model in distilling knowledge at the sequence model encoder. This work proposes a novel and competitive architecture for lip-reading, as we demonstrate a noticeable improvement in performance, setting a new benchmark equals to 88.64% on the LRW dataset.

📄 PDF Abstract BibTeX arXiv:2207.05692

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationLipreadingLip Readingspeech-recognitionSpeech RecognitionVisual Speech Recognition

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

Responsive Listening Head Generation: A Benchmark Dataset and Baseline

2021-12-27 · Mohan Zhou, Yalong Bai, Wei zhang, Ting Yao 외

We present a new listening head generation benchmark, for synthesizing responsive feedbacks of a listener (e.g., nod, smile) during a face-to-face conversation. As the indispensable complement to talking heads generation…

Talking Head GenerationTranslation

The ICASSP SP Cadenza Challenge: Music Demixing/Remixing for Hearing Aids

2023-10-05 · Gerardo Roa Dabike, Michael A. Akeroyd, Scott Bannister, Jon Barker 외

This paper reports on the design and results of the 2024 ICASSP SP Cadenza Challenge: Music Demixing/Remixing for Hearing Aids. The Cadenza project is working to enhance the audio quality of music for those with a hearin…

Generalization for slowly mixing processes

2023-04-28 · Andreas Maurer

A bound uniform over various loss-classes is given for data generated by stationary and phi-mixing processes, where the mixing time (the time needed to obtain approximate independence) enters the sample complexity only i…

Separating form and meaning: Using self-consistency to quantify task understanding across multiple senses

2023-05-19 · Xenia Ohmer, Elia Bruni, Dieuwke Hupkes

At the staggering pace with which the capabilities of large language models (LLMs) are increasing, creating future-proof evaluation sets to assess their understanding becomes more and more challenging. In this paper, we …

BenchmarkingForm

Deep generative demixing: Recovering Lipschitz signals from noisy subgaussian mixtures

2020-10-13 · Aaron Berk

Generative neural networks (GNNs) have gained renown for efficaciously capturing intrinsic low-dimensional structure in natural images. Here, we investigate the subgaussian demixing problem for two Lipschitz signals, wit…

compressed sensing