paper-with-me

홈 › Papers

Understanding the visual speech signal

2017-10-03 · Helen L. Bear

For machines to lipread, or understand speech from lip movement, they decode lip-motions (known as visemes) into the spoken sounds. We investigate the visual speech channel to further our understanding of visemes. This has applications beyond machine lipreading; speech therapists, animators, and psychologists can benefit from this work. We explain the influence of speaker individuality, and demonstrate how one can use visemes to boost lipreading.

📄 PDF Abstract BibTeX arXiv:1710.01351

Code (0)

등록된 구현이 없습니다.

Tasks

Lipreading

Similar Papers 제목 키워드 기반

"Notic My Speech" -- Blending Speech Patterns With Multimedia

2020-06-12 · Dhruva Sahrawat, Yaman Kumar, Shashwat Aggarwal, Yifang Yin 외

Speech as a natural signal is composed of three parts - visemes (visual part of speech), phonemes (spoken part of speech), and language (the imposed structure). However, video as a medium for the delivery of speech and a…

speech-recognitionSpeech RecognitionVideo CompressionVisual Speech Recognition

SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models

2026-05-31 · Chenshuang Zhang, Kyeong Seon Kim, Chengxin Liu, Tae-Hyun Oh arxiv

Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination. Existing benchmarks focus on environmental sounds (e.g., dog barking) to indicate…

Audio-Visual Speech Separation Using Cross-Modal Correspondence Loss

2021-03-02 · Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka 외

We present an audio-visual speech separation learning method that considers the correspondence between the separated signals and the visual signals to reflect the speech characteristics during training. Audio-visual spee…

Speech Separation

SViQA: A Unified Speech-Vision Multimodal Model for Textless Visual Question Answering

2025-04-01 · Bingxin Li

Multimodal models integrating speech and vision hold significant potential for advancing human-computer interaction, particularly in Speech-Based Visual Question Answering (SBVQA) where spoken questions about images requ…

cross-modal alignmentQuestion AnsweringVisual Question Answering

SynesLM: A Unified Approach for Audio-visual Speech Recognition and Translation via Language Model and Synthetic Data

2024-08-01 · Yichen Lu, Jiaqi Song, Xuankai Chang, Hengwei Bian 외

In this work, we present SynesLM, an unified model which can perform three multimodal language understanding tasks: audio-visual automatic speech recognition(AV-ASR) and visual-aided speech/machine translation(VST/VMT). …

Audio-Visual Speech RecognitionAutomatic Speech RecognitionLanguage ModelingLanguage Modelling+4