paper-with-me

홈 › Papers

Alternative Visual Units for an Optimized Phoneme-Based Lipreading System

2019-09-16

Lipreading is understanding speech from observed lip movements. An observed series of lip motions is an ordered sequence of visual lip gestures. These gestures are commonly known, but as yet are not formally defined, as `visemes'. In this article, we describe a structured approach which allows us to create speaker-dependent visemes with a fixed number of visemes within each set. We create sets of visemes for sizes two to 45. Each set of visemes is based upon clustering phonemes, thus each set has a unique phoneme-to-viseme mapping. We first present an experiment using these maps and the Resource Management Audio-Visual (RMAV) dataset which shows the effect of changing the viseme map size in speaker-dependent machine lipreading and demonstrate that word recognition with phoneme classifiers is possible. Furthermore, we show that there are intermediate units between visemes and phonemes which are better still. Second, we present a novel two-pass training scheme for phoneme classifiers. This approach uses our new intermediary visual units from our first experiment in the first pass as classifiers; before using the phoneme-to-viseme maps, we retrain these into phoneme classifiers. This method significantly improves on previous lipreading results with RMAV speakers.

📄 PDF Abstract BibTeX arXiv:1909.07147

Code (0)

등록된 구현이 없습니다.

Tasks

LipreadingManagement

Similar Papers 제목 키워드 기반

Comparing phonemes and visemes with DNN-based lipreading

2018-05-08 · Kwanchiva Thangthai, Helen L. Bear, Richard Harvey

There is debate if phoneme or viseme units are the most effective for a lipreading system. Some studies use phoneme units even though phonemes describe unique short sounds; other studies tried to improve lipreading accur…

DecoderLipreading

Visual Speech Language Models

2018-09-14

Language models (LM) are very powerful in lipreading systems. Language models built upon the ground truth utterances of datasets learn grammar and structure rules of words and sentences (the latter in the case of continu…

Language ModelingLanguage ModellingLipreading

Decoding visemes: improving machine lipreading

2017-10-03 · Helen L. Bear

Machine lipreading (MLR) is speech recognition from visual cues and a niche research problem in speech processing & computer vision. Current challenges fall into two groups: the content of the video, such as rate of spee…

ClusteringGeneral ClassificationLipreadingRobust classification+2

Visual speech recognition: aligning terminologies for better understanding

2017-10-03 · Helen L. Bear, Sarah Taylor

We are at an exciting time for machine lipreading. Traditional research stemmed from the adaptation of audio recognition systems. But now, the computer vision community is also participating. This joining of two previous…

Lipreadingspeech-recognitionSpeech RecognitionVisual Speech Recognition

The Lipreading Gap: Do VSR Models Perceive Visual Speech Like Human Lipreaders?

2026-06-05 · Rishabh Jain, Naomi Harte arxiv

Visual speech recognition (VSR) models now surpass human lipreaders on benchmarks, but do such gains establish human-like visual speech perception? To explore this, we compare three VSR systems with human baselines on th…

Visual Speech Recognition