paper-with-me

홈 › Papers

Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip Reading

2022-04-04 · The AAAI Conference on Artificial Intelligence (AAAI) 2022 3 · Minsu Kim, Jeong Hun Yeo, Yong Man Ro

Recognizing speech from silent lip movement, which is called lip reading, is a challenging task due to 1) the inherent information insufficiency of lip movement to fully represent the speech, and 2) the existence of homophenes that have similar lip movement with different pronunciations. In this paper, we try to alleviate the aforementioned two challenges in lip reading by proposing a Multi-head Visual-audio Memory (MVM). Firstly, MVM is trained with audio-visual datasets and remembers audio representations by modelling the inter-relationships of paired audio-visual representations. At the inference stage, visual input alone can extract the saved audio representation from the memory by examining the learned inter-relationships. Therefore, the lip reading model can complement the insufficient visual information with the extracted audio representations. Secondly, MVM is composed of multi-head key memories for saving visual features and one value memory for saving audio knowledge, which is designed to distinguish the homophenes. With the multi-head key memories, MVM extracts possible candidate audio features from the memory, which allows the lip reading model to consider the possibility of which pronunciations can be represented from the input lip movement. This also can be viewed as an explicit implementation of the one-to-many mapping of viseme-to-phoneme. Moreover, MVM is employed in multi-temporal levels to consider the context when retrieving the memory and distinguish the homophenes. Extensive experimental results verify the effectiveness of the proposed method in lip reading and in distinguishing the homophenes.

📄 PDF Abstract BibTeX arXiv:2204.01725

Code (1)

ms-dot-k/Multi-head-Visual-Audio-Memory pytorch

Tasks

LipreadingLip Reading

Similar Papers 제목 키워드 기반

SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization

2024-06-18 · Young Jin Ahn, Jungwoo Park, Sangha Park, Jonghyun Choi 외

Visual Speech Recognition (VSR) stands at the intersection of computer vision and speech recognition, aiming to interpret spoken content from visual cues. A prominent challenge in VSR is the presence of homophenes-visual…

Landmark-based LipreadingLipreadingspeech-recognitionSpeech Recognition+1

A Comprehensive Multi-scale Approach for Speech and Dynamics Synchrony in Talking Head Generation

2023-07-04 · Louis Airale, Dominique Vaufreydaz, Xavier Alameda-Pineda

Animating still face images with deep generative models using a speech input signal is an active research topic and has seen important recent progress.However, much of the effort has been put into lip syncing and renderi…

Talking Head Generation

Let There Be Sound: Reconstructing High Quality Speech from Silent Videos

2023-08-29 · Ji-Hoon Kim, Jaehun Kim, Joon Son Chung

The goal of this work is to reconstruct high quality speech from lip motions alone, a task also known as lip-to-speech. A key challenge of lip-to-speech systems is the one-to-many mapping caused by (1) the existence of h…

AV-Gaze: A Study on the Effectiveness of Audio Guided Visual Attention Estimation for Non-Profilic Faces

2022-07-07 · Shreya Ghosh, Abhinav Dhall, Munawar Hayat, Jarrod Knibbe

In challenging real-life conditions such as extreme head-pose, occlusions, and low-resolution images where the visual information fails to estimate visual attention/gaze direction, audio signals could provide important a…

Improving Visual Speech Enhancement Network by Learning Audio-visual Affinity with Multi-head Attention

2022-06-30 · Xinmeng Xu, Yang Wang, Jie Jia, Binbin Chen 외

Audio-visual speech enhancement system is regarded as one of promising solutions for isolating and enhancing speech of desired speaker. Typical methods focus on predicting clean speech spectrum via a naive convolution ne…

DecoderSpeech Enhancement