paper-with-me

홈 › Papers

Visual Speech Language Models

2018-09-14

Language models (LM) are very powerful in lipreading systems. Language models built upon the ground truth utterances of datasets learn grammar and structure rules of words and sentences (the latter in the case of continuous speech). However, visual co-articulation effects in visual speech signals damage the performance of visual speech LM's as visually, people do not utter what the language model expects. These models are commonplace but while higher-order N-gram LM's may improve classification rates, the cost of this model is disproportionate to the common goal of developing more accurate classifiers. So we compare which unit would best optimize a lipreading (visual speech) LM to observe their limitations. We compare three units; visemes (visual speech units) \cite{lan2010improving}, phonemes (audible speech units), and words.

📄 PDF Abstract BibTeX arXiv:1809.06800

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLipreading

Similar Papers 제목 키워드 기반

AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation

2023-12-05 · CVPR 2024 1 · Jeongsoo Choi, Se Jin Park, Minsu Kim, Yong Man Ro

This paper proposes a novel direct Audio-Visual Speech to Audio-Visual Speech Translation (AV2AV) framework, where the input and output of the system are multimodal (i.e., audio and visual speech). With the proposed AV2A…

Self-Supervised LearningSpeech-to-Speech TranslationTranslation

Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations

2025-03-08 · Jeong Hun Yeo, Minsu Kim, Chae Won Kim, Stavros Petridis 외

We explore a novel zero-shot Audio-Visual Speech Recognition (AVSR) framework, dubbed Zero-AVSR, which enables speech recognition in target languages without requiring any audio-visual speech data in those languages. Spe…

Audio-Visual Speech RecognitionMulti-Task Learningspeech-recognitionSpeech Recognition+1

MixSpeech: Cross-Modality Self-Learning with Audio-Visual Stream Mixup for Visual Speech Translation and Recognition

2023-03-09 · ICCV 2023 1 · Xize Cheng, Linjun Li, Tao Jin, Rongjie Huang 외

Multi-media communications facilitate global interaction among people. However, despite researchers exploring cross-lingual translation techniques such as machine translation and audio speech translation to overcome lang…

Lip ReadingMachine TranslationSelf-LearningTransfer Learning+2

Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech Representation

2024-01-18 · Minsu Kim, Jeong Hun Yeo, Se Jin Park, Hyeongseop Rha 외

This paper explores sentence-level multilingual Visual Speech Recognition (VSR) that can recognize different languages with a single trained model. As the massive multilingual modeling of visual data requires huge comput…

Sentencespeech-recognitionSpeech RecognitionVisual Speech Recognition

CSLNSpeech: solving extended speech separation problem with the help of Chinese sign language

2020-07-21 · Jiasong Wu, Xuan Li, Taotao Li, Fanman Meng 외

Previous audio-visual speech separation methods use the synchronization of the speaker's facial movement and speech in the video to supervise the speech separation in a self-supervised way. In this paper, we propose a mo…

Self-Supervised LearningSpeech Separation