paper-with-me

Lip Reading

3개 벤치마크 · 논문 158편 · 이 태스크의 논문 보기 →

Benchmarks

LRW

결과 1개

Most implemented

Deep Audio-Visual Speech Recognition

2018-09-06 · 구현 4개

Papers

TVTA: Trajectory-Aware Viseme-Guided Temporal Aggregation for Event-Based Lip Reading

2026-07-09 · Jingrong Zheng, Hongwei Ren, Xiangqian Wu arxiv

Event-based lip reading has recently emerged as a promising direction for visual speech recognition, benefiting from the high temporal resolution and motion sensitivity of event cameras. However, existing methods typical…

Visual Speech RecognitionLip Reading

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness

2026-07-09 · Iulia-Maria Udrea, Alexandra Diaconu, Bogdan Alexe arxiv

We introduce VSRo-200, the first large-scale dataset for visual speech recognition (lip reading) in Romanian, comprising 200 hours of real-world podcast videos. All samples are annotated with pseudo-labels generated by a…

Audio-Visual Speech RecognitionDomain GeneralizationLip Reading

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

2025-12-15 · Foivos Paraperas Papantoniou, Stathis Galanakis, Rolandos Alexandros Potamias, Bernhard Kainz 외 arxiv

This paper presents STARCaster, an identity-aware spatio-temporal video diffusion model that addresses both speech-driven portrait animation and free-viewpoint talking portrait synthesis, given an identity embedding or r…

Lip Reading

GLip: A Global-Local Integrated Progressive Framework for Robust Visual Speech Recognition

2025-09-19 · Tianyue Wang, Shuang Yang, Shiguang Shan, Xilin Chen arxiv

Visual speech recognition (VSR), also known as lip reading, is the task of recognizing speech from silent video. Despite significant advancements in VSR over recent decades, most existing methods pay limited attention to…

Visual Speech RecognitionLip Reading

Towards Inclusive Communication: A Unified Framework for Generating Spoken Language from Sign, Lip, and Audio

2025-08-28 · Jeong Hun Yeo, Hyeongseop Rha, Sungjune Park, Junil Won 외 arxiv

Audio is the primary modality for human communication and has driven the success of Automatic Speech Recognition (ASR) technologies. However, such audio-centric systems inherently exclude individuals who are deaf or hard…

Audio-Visual Speech RecognitionSign Language TranslationText GenerationLip Reading

VisualSpeaker: Visually-Guided 3D Avatar Lip Synthesis

2025-07-08 · Alexandre Symeonidis-Herzig, Özge Mercanoğlu Sincan, Richard Bowden

Realistic, high-fidelity 3D facial animations are crucial for expressive avatar systems in human-computer interaction and accessibility. Although prior methods show promising quality, their reliance on the mesh domain li…

Automatic Speech RecognitionLip Readingspeech-recognitionSpeech Recognition+1

전체 158편 보기 →