Lip Reading
3개 벤치마크 · 논문 158편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Deep Audio-Visual Speech Recognition
Combining Residual Networks with LSTMs for Lipreading
VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness
End-to-end Audio-visual Speech Recognition with Conformers
Papers
TVTA: Trajectory-Aware Viseme-Guided Temporal Aggregation for Event-Based Lip Reading
Event-based lip reading has recently emerged as a promising direction for visual speech recognition, benefiting from the high temporal resolution and motion sensitivity of event cameras. However, existing methods typical…
Visual Speech RecognitionLip ReadingVSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness
We introduce VSRo-200, the first large-scale dataset for visual speech recognition (lip reading) in Romanian, comprising 200 hours of real-world podcast videos. All samples are annotated with pseudo-labels generated by a…
Audio-Visual Speech RecognitionDomain GeneralizationLip ReadingSTARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
This paper presents STARCaster, an identity-aware spatio-temporal video diffusion model that addresses both speech-driven portrait animation and free-viewpoint talking portrait synthesis, given an identity embedding or r…
Lip ReadingGLip: A Global-Local Integrated Progressive Framework for Robust Visual Speech Recognition
Visual speech recognition (VSR), also known as lip reading, is the task of recognizing speech from silent video. Despite significant advancements in VSR over recent decades, most existing methods pay limited attention to…
Visual Speech RecognitionLip ReadingTowards Inclusive Communication: A Unified Framework for Generating Spoken Language from Sign, Lip, and Audio
Audio is the primary modality for human communication and has driven the success of Automatic Speech Recognition (ASR) technologies. However, such audio-centric systems inherently exclude individuals who are deaf or hard…
Audio-Visual Speech RecognitionSign Language TranslationText GenerationLip ReadingVisualSpeaker: Visually-Guided 3D Avatar Lip Synthesis
Realistic, high-fidelity 3D facial animations are crucial for expressive avatar systems in human-computer interaction and accessibility. Although prior methods show promising quality, their reliance on the mesh domain li…
Automatic Speech RecognitionLip Readingspeech-recognitionSpeech Recognition+1