Landmark-based Lipreading
2개 벤치마크 · 논문 4편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Papers
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
Visual Speech Recognition (VSR) stands at the intersection of computer vision and speech recognition, aiming to interpret spoken content from visual cues. A prominent challenge in VSR is the presence of homophenes-visual…
Landmark-based LipreadingLipreadingspeech-recognitionSpeech Recognition+1Another Point of View on Visual Speech Recognition
Standard Visual Speech Recognition (VSR) systems directly process images as input features without any apriori link between raw pixel data and facial traits. Pixel information is smartly sieved when facial landmarks are …
Landmark-based Lipreadingspeech-recognitionSpeech RecognitionVisual Speech RecognitionAdaptive Semantic-Spatio-Temporal Graph Convolutional Network for Lip Reading
The goal of this work is to recognize words, phrases, and sentences being spoken by a talking face without given the audio. Current deep learning approaches for lip reading focus on exploring the appearance and optical f…
Landmark-based LipreadingLip ReadingOptical Flow EstimationLip Graph Assisted Audio-Visual Speech Recognition Using Bidirectional Synchronous Fusion
Current studies have shown that extracting representative visual features and efficiently fusing audio and visual modalities are vital for audio-visual speech recognition (AVSR), but these are still challenging. To this …
Audio-Visual Speech RecognitionLandmark-based Lipreadingspeech-recognitionSpeech Recognition+1