paper-with-me

Papers

Decoding visemes: improving machine lipreading

2017-10-03 · Helen L. Bear

Machine lipreading (MLR) is speech recognition from visual cues and a niche research problem in speech processing & computer vision. Current challenges fall into two groups: the content of the video, such as rate of speech or; the parameters of the video recording e.g, video resolution. We show that HD video is not needed to successfully lipread with a computer. The term "viseme" is used in machine lipreading to represent a visual cue or gesture which corresponds to a subgroup of phonemes where the phonemes are visually indistinguishable. A phoneme is the smallest sound one can utter, because there are more phonemes per viseme, maps between units show a many-to-one relationship. Many maps have been presented, we compare these and our results show Lee's is best. We propose a new method of speaker-dependent phoneme-to-viseme maps and compare these to Lee's. Our results show the sensitivity of phoneme clustering and we use our new knowledge to augment a conventional MLR system. It has been observed in MLR, that classifiers need training on test subjects to achieve accuracy. Thus machine lipreading is highly speaker-dependent. Conversely speaker independence is robust classification of non-training speakers. We investigate the dependence of phoneme-to-viseme maps between speakers and show there is not a high variability of visemes, but there is high variability in trajectory between visemes of individual speakers with the same ground truth. This implies a dependency upon the number of visemes within each set for each individual. We show that prior phoneme-to-viseme maps rarely have enough visemes and the optimal size, which varies by speaker, ranges from 11-35. Finally we decode from visemes back to phonemes and into words. Our novel approach uses the optimum range visemes within hierarchical training of phoneme classifiers and demonstrates a significant increase in classification accuracy.

📄 PDF Abstract BibTeX arXiv:1710.01288

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringGeneral ClassificationLipreadingRobust classificationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Understanding the visual speech signal

2017-10-03 · Helen L. Bear

For machines to lipread, or understand speech from lip movement, they decode lip-motions (known as visemes) into the spoken sounds. We investigate the visual speech channel to further our understanding of visemes. This h…

Lipreading

Decoding visemes: improving machine lipreading

2017-10-03 · Helen L. Bear, Richard Harvey

To undertake machine lip-reading, we try to recognise speech from a visual signal. Current work often uses viseme classification supported by language models with varying degrees of success. A few recent works suggest ph…

ClassificationGeneral ClassificationLipreadingLip Reading

Comparing phonemes and visemes with DNN-based lipreading

2018-05-08 · Kwanchiva Thangthai, Helen L. Bear, Richard Harvey

There is debate if phoneme or viseme units are the most effective for a lipreading system. Some studies use phoneme units even though phonemes describe unique short sounds; other studies tried to improve lipreading accur…

DecoderLipreading

Alternative Visual Units for an Optimized Phoneme-Based Lipreading System

2019-09-16

Lipreading is understanding speech from observed lip movements. An observed series of lip motions is an ordered sequence of visual lip gestures. These gestures are commonly known, but as yet are not formally defined, as …

LipreadingManagement

Visual gesture variability between talkers in continuous visual speech

2017-10-03 · Helen L. Bear

Recent adoption of deep learning methods to the field of machine lipreading research gives us two options to pursue to improve system performance. Either, we develop end-to-end systems holistically or, we experiment to f…

Lipreading