paper-with-me

홈 › Papers

Learning Separable Hidden Unit Contributions for Speaker-Adaptive Lip-Reading

2023-10-08 · Songtao Luo, Shuang Yang, Shiguang Shan, Xilin Chen

In this paper, we propose a novel method for speaker adaptation in lip reading, motivated by two observations. Firstly, a speaker's own characteristics can always be portrayed well by his/her few facial images or even a single image with shallow networks, while the fine-grained dynamic features associated with speech content expressed by the talking face always need deep sequential networks to represent accurately. Therefore, we treat the shallow and deep layers differently for speaker adaptive lip reading. Secondly, we observe that a speaker's unique characteristics ( e.g. prominent oral cavity and mandible) have varied effects on lip reading performance for different words and pronunciations, necessitating adaptive enhancement or suppression of the features for robust lip reading. Based on these two observations, we propose to take advantage of the speaker's own characteristics to automatically learn separable hidden unit contributions with different targets for shallow layers and deep layers respectively. For shallow layers where features related to the speaker's characteristics are stronger than the speech content related features, we introduce speaker-adaptive features to learn for enhancing the speech content features. For deep layers where both the speaker's features and the speech content features are all expressed well, we introduce the speaker-adaptive features to learn for suppressing the speech content irrelevant noise for robust lip reading. Our approach consistently outperforms existing methods, as confirmed by comprehensive analysis and comparison across different settings. Besides the evaluation on the popular LRW-ID and GRID datasets, we also release a new dataset for evaluation, CAS-VSR-S68h, to further assess the performance in an extreme setting where just a few speakers are available but the speech content covers a large and diversified range.

📄 PDF Abstract BibTeX arXiv:2310.05058

Code (1)

jinchiniao/LSHUC 공식 구현 pytorch

Tasks

Lip Reading

Similar Papers 제목 키워드 기반

Learning Hidden Unit Contributions for Unsupervised Acoustic Model Adaptation

2016-01-12 · Pawel Swietojanski, Jinyu Li, Steve Renals

This work presents a broad study on the adaptation of neural network acoustic models by means of learning hidden unit contributions (LHUC) -- a method that linearly re-combines hidden units in a speaker- or environment-d…

speech-recognitionSpeech Recognition

Investigation of Data Augmentation Techniques for Disordered Speech Recognition

2022-01-14 · Mengzhe Geng, Xurong Xie, Shansong Liu, Jianwei Yu 외

Disordered speech recognition is a highly challenging task. The underlying neuro-motor conditions of people with speech disorders, often compounded with co-occurring physical disabilities, lead to the difficulty in colle…

Data Augmentationspeech-recognitionSpeech Recognition

Confidence Score Based Conformer Speaker Adaptation for Speech Recognition

2022-06-24 · Jiajun Deng, Xurong Xie, Tianzi Wang, Mingyu Cui 외

A key challenge for automatic speech recognition (ASR) systems is to model the speaker level variability. In this paper, compact speaker dependent learning hidden unit contributions (LHUC) are used to facilitate both spe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeech-recognition+1

Bayesian Learning for Deep Neural Network Adaptation

2020-12-14 · Xurong Xie, Xunying Liu, Tan Lee, Lan Wang

A key task for speech recognition systems is to reduce the mismatch between training and evaluation data that is often attributable to speaker differences. Speaker adaptation techniques play a vital role to reduce the mi…

speech-recognitionSpeech RecognitionVariational Inference

Unsupervised Model-based speaker adaptation of end-to-end lattice-free MMI model for speech recognition

2022-11-17 · Xurong Xie, Xunying Liu, Hui Chen, Hongan Wang

Modeling the speaker variability is a key challenge for automatic speech recognition (ASR) systems. In this paper, the learning hidden unit contributions (LHUC) based adaptation techniques with compact speaker dependent …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)modelspeech-recognition+1