paper-with-me

Visual Speech Recognition

2개 벤치마크 · 논문 198편 · 이 태스크의 논문 보기 →

Benchmarks

LRS2

결과 6개

LRS3-TED

결과 6개

Most implemented

Deep Audio-Visual Speech Recognition

2018-09-06 · 구현 4개

Papers

TVTA: Trajectory-Aware Viseme-Guided Temporal Aggregation for Event-Based Lip Reading

2026-07-09 · Jingrong Zheng, Hongwei Ren, Xiangqian Wu arxiv

Event-based lip reading has recently emerged as a promising direction for visual speech recognition, benefiting from the high temporal resolution and motion sensitivity of event cameras. However, existing methods typical…

Visual Speech RecognitionLip Reading

The Lipreading Gap: Do VSR Models Perceive Visual Speech Like Human Lipreaders?

2026-06-05 · Rishabh Jain, Naomi Harte arxiv

Visual speech recognition (VSR) models now surpass human lipreaders on benchmarks, but do such gains establish human-like visual speech perception? To explore this, we compare three VSR systems with human baselines on th…

Visual Speech Recognition

Head-Pose-Aware Visual Speech Recognition with FiLM Modulation

2026-05-30 · Matthew Kit Khinn Teng, Haibo Zhang, Takeshi Saitoh arxiv

Visual Speech Recognition (VSR) aims to recognize speech from visual cues such as lip movements, but its performance is fundamentally limited by viseme ambiguity and pose-induced variations that introduce geometric disto…

Visual Speech Recognition

Diffusion Large Language Models for Visual Speech Recognition

2026-05-27 · Jeong Hun Yeo, Chae Won Kim, Hyeongseop Rha, Yong Man Ro arxiv

Existing Visual Speech Recognition (VSR) systems commonly rely on left-to-right autoregressive decoding, which can force premature decisions on visually ambiguous tokens before sufficient context is available. We propose…

Visual Speech Recognition

Cascade-Free Mandarin Visual Speech Recognition via Semantic-Guided Cross-Representation Alignment

2026-03-23 · Lei Yang, Yi He, Fei Wu, Shilin Wang arxiv

Chinese mandarin visual speech recognition (VSR) is a task that has advanced in recent years, yet still lags behind the performance on non-tonal languages such as English. One primary challenge arises from the tonal natu…

Visual Speech Recognition

Visual-Informed Speech Enhancement Using Attention-Based Beamforming

2026-03-05 · Chihyun Liu, Jiaxuan Fan, Mingtung Sun, Michael Anthony 외 arxiv

Recent studies have demonstrated that incorporating auxiliary information, such as speaker voiceprint or visual cues, can substantially improve Speech Enhancement (SE) performance. However, single-channel methods often y…

Visual Speech RecognitionSpeaker IdentificationSpeech EnhancementActivity Detection

전체 198편 보기 →