Understanding the visual speech signal
For machines to lipread, or understand speech from lip movement, they decode lip-motions (known as visemes) into the spoken sounds. We investigate the visual speech channel to further our understanding of visemes. This has applications beyond machine lipreading; speech therapists, animators, and psychologists can benefit from this work. We explain the influence of speaker individuality, and demonstrate how one can use visemes to boost lipreading.
Code (0)
등록된 구현이 없습니다.
Tasks
LipreadingSimilar Papers 제목 키워드 기반
"Notic My Speech" -- Blending Speech Patterns With Multimedia
Speech as a natural signal is composed of three parts - visemes (visual part of speech), phonemes (spoken part of speech), and language (the imposed structure). However, video as a medium for the delivery of speech and a…
speech-recognitionSpeech RecognitionVideo CompressionVisual Speech RecognitionSVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models
Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination. Existing benchmarks focus on environmental sounds (e.g., dog barking) to indicate…
Audio-Visual Speech Separation Using Cross-Modal Correspondence Loss
We present an audio-visual speech separation learning method that considers the correspondence between the separated signals and the visual signals to reflect the speech characteristics during training. Audio-visual spee…
Speech SeparationSViQA: A Unified Speech-Vision Multimodal Model for Textless Visual Question Answering
Multimodal models integrating speech and vision hold significant potential for advancing human-computer interaction, particularly in Speech-Based Visual Question Answering (SBVQA) where spoken questions about images requ…
cross-modal alignmentQuestion AnsweringVisual Question AnsweringSynesLM: A Unified Approach for Audio-visual Speech Recognition and Translation via Language Model and Synthetic Data
In this work, we present SynesLM, an unified model which can perform three multimodal language understanding tasks: audio-visual automatic speech recognition(AV-ASR) and visual-aided speech/machine translation(VST/VMT). …
Audio-Visual Speech RecognitionAutomatic Speech RecognitionLanguage ModelingLanguage Modelling+4