Papers Lip Reading
“Lip Reading” 태그가 달린 논문 158편 · 필터 해제
TVTA: Trajectory-Aware Viseme-Guided Temporal Aggregation for Event-Based Lip Reading
Event-based lip reading has recently emerged as a promising direction for visual speech recognition, benefiting from the high temporal resolution and motion sensitivity of event cameras. However, existing methods typical…
Visual Speech RecognitionLip ReadingVSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness
We introduce VSRo-200, the first large-scale dataset for visual speech recognition (lip reading) in Romanian, comprising 200 hours of real-world podcast videos. All samples are annotated with pseudo-labels generated by a…
Audio-Visual Speech RecognitionDomain GeneralizationLip ReadingSTARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
This paper presents STARCaster, an identity-aware spatio-temporal video diffusion model that addresses both speech-driven portrait animation and free-viewpoint talking portrait synthesis, given an identity embedding or r…
Lip ReadingGLip: A Global-Local Integrated Progressive Framework for Robust Visual Speech Recognition
Visual speech recognition (VSR), also known as lip reading, is the task of recognizing speech from silent video. Despite significant advancements in VSR over recent decades, most existing methods pay limited attention to…
Visual Speech RecognitionLip ReadingTowards Inclusive Communication: A Unified Framework for Generating Spoken Language from Sign, Lip, and Audio
Audio is the primary modality for human communication and has driven the success of Automatic Speech Recognition (ASR) technologies. However, such audio-centric systems inherently exclude individuals who are deaf or hard…
Audio-Visual Speech RecognitionSign Language TranslationText GenerationLip ReadingVisualSpeaker: Visually-Guided 3D Avatar Lip Synthesis
Realistic, high-fidelity 3D facial animations are crucial for expressive avatar systems in human-computer interaction and accessibility. Although prior methods show promising quality, their reliance on the mesh domain li…
Automatic Speech RecognitionLip Readingspeech-recognitionSpeech Recognition+1SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin Transformer
This paper presents an efficient visual speech encoder for lip reading. While most recent lip reading studies have been based on the ResNet architecture and have achieved significant success, they are not sufficiently su…
Audio-Visual Speech RecognitionLip ReadingSpeech Enhancementspeech-recognition+3Transforming faces into video stories -- VideoFace2.0
Face detection and face recognition have been in the focus of vision community since the very beginnings. Inspired by the success of the original Videoface digitizer, a pioneering device that allowed users to capture vid…
Face DetectionFace RecognitionLip Readingspeech-recognition+2Development and evaluation of a deep learning algorithm for German word recognition from lip movements
When reading lips, many people benefit from additional visual information from the lip movements of the speaker, which is, however, very error prone. Algorithms for lip reading with artificial intelligence based on artif…
Lip Readingspeech-recognitionSpeech RecognitionChinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides
Incorporating visual modalities to assist Automatic Speech Recognition (ASR) tasks has led to significant improvements. However, existing Audio-Visual Speech Recognition (AVSR) datasets and methods typically rely solely …
Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Lip Reading+3VALLR: Visual ASR Language Model for Lip Reading
Lip Reading, or Visual Automatic Speech Recognition (V-ASR), is a complex task requiring the interpretation of spoken language exclusively from visual cues, primarily lip movements and facial expressions. This task is es…
Automatic Speech RecognitionLanguage ModelingLanguage ModellingLarge Language Model+3Lend a Hand: Semi Training-Free Cued Speech Recognition via MLLM-Driven Hand Modeling for Barrier-free Communication
Cued Speech (CS) is an innovative visual communication system that integrates lip-reading with hand coding, designed to enhance effective communication for individuals with hearing impairments. Automatic CS Recognition (…
Lip ReadingPrompt Engineeringspeech-recognitionSpeech Recognition+1Integrating Persian Lip Reading in Surena-V Humanoid Robot for Human-Robot Interaction
Lip reading is vital for robots in social settings, improving their ability to understand human communication. This skill allows them to communicate more easily in crowded environments, especially in caregiving and custo…
Landmark TrackingLip Readingspeech-recognitionSpeech RecognitionGLaM-Sign: Greek Language Multimodal Lip Reading with Integrated Sign Language Accessibility
The Greek Language Multimodal Lip Reading with Integrated Sign Language Accessibility (GLaM-Sign) [1] is a groundbreaking resource in accessibility and multimodal AI, designed to support Deaf and Hard-of-Hearing (DHH) in…
Lip ReadingSign Language TranslationLipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
Visual speech recognition (VSR), commonly known as lip reading, has garnered significant attention due to its wide-ranging practical applications. The advent of deep learning techniques and advancements in hardware capab…
Lip Readingspeech-recognitionSpeech RecognitionVideo Generation+1Spatio-temporal Transformers for Action Unit Classification with Event Cameras
Face analysis has been studied from different angles to infer emotion, poses, shapes, and landmarks. Traditionally RGB cameras are used, yet for fine-grained tasks standard sensors might not be up to the task due to thei…
Lip ReadingQuantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective
In the field of spoken language processing, audio-visual speech processing is receiving increasing research attention. Key components of this research include tasks such as lip reading, audio-visual speech recognition, a…
Audio-Visual Speech RecognitionLip Readingspeech-recognitionSpeech Recognition+2Neuromorphic Facial Analysis with Cross-Modal Supervision
Traditional approaches for analyzing RGB frames are capable of providing a fine-grained understanding of a face from different angles by inferring emotions, poses, shapes, landmarks. However, when it comes to subtle move…
Lip ReadingRAL:Redundancy-Aware Lipreading Model Based on Differential Learning with Symmetric Views
Lip reading involves interpreting a speaker's speech by analyzing sequences of lip movements. Currently, most models regard the left and right halves of the lips as a symmetrical whole, lacking a thorough investigation o…
LipreadingLip ReadingPersonalized Lip Reading: Adapting to Your Unique Lip Movements with Vision and Language
Lip reading aims to predict spoken language by analyzing lip movements. Despite advancements in lip reading technologies, performance degrades when models are applied to unseen speakers due to their sensitivity to variat…
Lip ReadingSentence