Visual Speech Recognition
2개 벤치마크 · 논문 198편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Deep Audio-Visual Speech Recognition
Combining Residual Networks with LSTMs for Lipreading
End-to-end Audio-visual Speech Recognition with Conformers
The NPU-ASLP-LiAuto System Description for Visual Speech Recognition in CNVSRC 2023
Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels
Papers
TVTA: Trajectory-Aware Viseme-Guided Temporal Aggregation for Event-Based Lip Reading
Event-based lip reading has recently emerged as a promising direction for visual speech recognition, benefiting from the high temporal resolution and motion sensitivity of event cameras. However, existing methods typical…
Visual Speech RecognitionLip ReadingThe Lipreading Gap: Do VSR Models Perceive Visual Speech Like Human Lipreaders?
Visual speech recognition (VSR) models now surpass human lipreaders on benchmarks, but do such gains establish human-like visual speech perception? To explore this, we compare three VSR systems with human baselines on th…
Visual Speech RecognitionHead-Pose-Aware Visual Speech Recognition with FiLM Modulation
Visual Speech Recognition (VSR) aims to recognize speech from visual cues such as lip movements, but its performance is fundamentally limited by viseme ambiguity and pose-induced variations that introduce geometric disto…
Visual Speech RecognitionDiffusion Large Language Models for Visual Speech Recognition
Existing Visual Speech Recognition (VSR) systems commonly rely on left-to-right autoregressive decoding, which can force premature decisions on visually ambiguous tokens before sufficient context is available. We propose…
Visual Speech RecognitionCascade-Free Mandarin Visual Speech Recognition via Semantic-Guided Cross-Representation Alignment
Chinese mandarin visual speech recognition (VSR) is a task that has advanced in recent years, yet still lags behind the performance on non-tonal languages such as English. One primary challenge arises from the tonal natu…
Visual Speech RecognitionVisual-Informed Speech Enhancement Using Attention-Based Beamforming
Recent studies have demonstrated that incorporating auxiliary information, such as speaker voiceprint or visual cues, can substantially improve Speech Enhancement (SE) performance. However, single-channel methods often y…
Visual Speech RecognitionSpeaker IdentificationSpeech EnhancementActivity Detection