paper-with-me

홈 › Papers

Which phoneme-to-viseme maps best improve visual-only computer lip-reading?

2017-10-03 · Helen L. Bear, Richard W. Harvey, Barry-John Theobald, Yuxuan Lan

A critical assumption of all current visual speech recognition systems is that there are visual speech units called visemes which can be mapped to units of acoustic speech, the phonemes. Despite there being a number of published maps it is infrequent to see the effectiveness of these tested, particularly on visual-only lip-reading (many works use audio-visual speech). Here we examine 120 mappings and consider if any are stable across talkers. We show a method for devising maps based on phoneme confusions from an automated lip-reading system, and we present new mappings that show improvements for individual talkers.

📄 PDF Abstract BibTeX arXiv:1710.01093

Code (0)

등록된 구현이 없습니다.

Tasks

Lip Readingspeech-recognitionSpeech RecognitionVisual Speech Recognition

Similar Papers 제목 키워드 기반

Decoding visemes: improving machine lipreading

2017-10-03 · Helen L. Bear

Machine lipreading (MLR) is speech recognition from visual cues and a niche research problem in speech processing & computer vision. Current challenges fall into two groups: the content of the video, such as rate of spee…

ClusteringGeneral ClassificationLipreadingRobust classification+2

Alternative Visual Units for an Optimized Phoneme-Based Lipreading System

2019-09-16

Lipreading is understanding speech from observed lip movements. An observed series of lip motions is an ordered sequence of visual lip gestures. These gestures are commonly known, but as yet are not formally defined, as …

LipreadingManagement

Finding phonemes: improving machine lip-reading

2017-10-03 · Helen L. Bear, Richard W. Harvey, Yuxuan Lan

In machine lip-reading there is continued debate and research around the correct classes to be used for recognition. In this paper we use a structured approach for devising speaker-dependent viseme classes, which enables…

Lip ReadingPhoneme Recognition

Phoneme-to-viseme mappings: the good, the bad, and the ugly

2018-05-08 · Helen L. Bear, Richard Harvey

Visemes are the visual equivalent of phonemes. Although not precisely defined, a working definition of a viseme is "a set of phonemes which have identical appearance on the lips". Therefore a phoneme falls into one visem…

Comparing heterogeneous visual gestures for measuring the diversity of visual speech signals

2018-05-08 · Helen L. Bear, Richard Harvey

Visual lip gestures observed whilst lipreading have a few working definitions, the most common two are; `the visual equivalent of a phoneme' and `phonemes which are indistinguishable on the lips'. To date there is no for…

ClusteringDiversityLipreading