paper-with-me

Papers

Transfer Learning from Visual Speech Recognition to Mouthing Recognition in German Sign Language

2025-05-20 · Dinh Nam Pham, Eleftherios Avramidis

Sign Language Recognition (SLR) systems primarily focus on manual gestures, but non-manual features such as mouth movements, specifically mouthing, provide valuable linguistic information. This work directly classifies mouthing instances to their corresponding words in the spoken language while exploring the potential of transfer learning from Visual Speech Recognition (VSR) to mouthing recognition in German Sign Language. We leverage three VSR datasets: one in English, one in German with unrelated words and one in German containing the same target words as the mouthing dataset, to investigate the impact of task similarity in this setting. Our results demonstrate that multi-task learning improves both mouthing recognition and VSR accuracy as well as model robustness, suggesting that mouthing recognition should be treated as a distinct but related task to VSR. This research contributes to the field of SLR by proposing knowledge transfer from VSR to SLR datasets with limited mouthing annotations.

📄 PDF Abstract BibTeX arXiv:2505.13784

Code (1)

nphamdinh/transfer-learning-vsr-mouthing-sign-language 공식 구현 pytorch

Tasks

Multi-Task LearningSign Language Recognitionspeech-recognitionSpeech RecognitionTransfer LearningVisual Speech Recognition

Methods 이 논문이 사용한 방법론

Focus 설명 없음
SLR Please enter a description about the method here

Similar Papers 제목 키워드 기반

VisualSpeaker: Visually-Guided 3D Avatar Lip Synthesis

2025-07-08 · Alexandre Symeonidis-Herzig, Özge Mercanoğlu Sincan, Richard Bowden

Realistic, high-fidelity 3D facial animations are crucial for expressive avatar systems in human-computer interaction and accessibility. Although prior methods show promising quality, their reliance on the mesh domain li…

Automatic Speech RecognitionLip Readingspeech-recognitionSpeech Recognition+1

Mouthing Recognition with OpenPose in Sign Language

2022-06-01 · SLTAT (LREC) 2022 6 · Maria Del Carmen Saenz

Many avatars focus on the hands and how they express sign language. However, sign language also uses mouth and face gestures to modify verbs, adjectives, or adverbs; these are known as non-manual components of the sign. …

Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech Recognition

2023-06-18 · Yuchen Hu, Ruizhe Li, Chen Chen, Chengwei Qin 외

Audio-visual speech recognition (AVSR) provides a promising solution to ameliorate the noise-robustness of audio-only speech recognition with visual information. However, most existing efforts still focus on audio modali…

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

BSL-1K: Scaling up co-articulated sign language recognition using mouthing cues

2020-07-23 · ECCV 2020 8 · Samuel Albanie, Gül Varol, Liliane Momeni, Triantafyllos Afouras 외

Recent progress in fine-grained gesture and action classification, and machine translation, point to the possibility of automated sign language recognition becoming a reality. A key stumbling block in making progress tow…

Action ClassificationKeyword SpottingMachine TranslationSign Language Recognition

Emotion Recognition in Speech using Cross-Modal Transfer in the Wild

2018-08-16 · Samuel Albanie, Arsha Nagrani, Andrea Vedaldi, Andrew Zisserman

Obtaining large, human labelled speech datasets to train models for emotion recognition is a notoriously challenging task, hindered by annotation cost and label ambiguity. In this work, we consider the task of learning e…

Emotion RecognitionFacial Emotion RecognitionFacial Expression Recognition (FER)Speech Emotion Recognition