Comparing heterogeneous visual gestures for measuring the diversity of visual speech signals
Visual lip gestures observed whilst lipreading have a few working
definitions, the most common two are; the visual equivalent of a phoneme' and
phonemes which are indistinguishable on the lips'. To date there is no formal
definition, in part because to date we have not established a two-way
relationship or mapping between visemes and phonemes. Some evidence suggests
that visual speech is highly dependent upon the speaker. So here, we use a
phoneme-clustering method to form new phoneme-to-viseme maps for both
individual and multiple speakers. We test these phoneme to viseme maps to
examine how similarly speakers talk visually and we use signed rank tests to
measure the distance between individuals. We conclude that broadly speaking,
speakers have the same repertoire of mouth gestures, where they differ is in
the use of the gestures.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringDiversityLipreadingSimilar Papers 제목 키워드 기반
Interpretable Diversity Analysis: Visualizing Feature Representations In Low-Cost Ensembles
Diversity is an important consideration in the construction of robust neural network ensembles. A collection of well trained models will generalize better if they are diverse in the patterns they respond to and the predi…
DiversityVisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning
The effectiveness of multimodal instruction tuning depends not only on dataset scale, but critically on whether training samples genuinely require visual reasoning. However, existing instruction datasets often contain a …
Visual ReasoningEMOTION: Expressive Motion Sequence Generation for Humanoid Robots with In-Context Learning
This paper introduces a framework, called EMOTION, for generating expressive motion sequences in humanoid robots, enhancing their ability to engage in humanlike non-verbal communication. Non-verbal cues such as facial ex…
DiversityIn-Context LearningSignCol: Open-Source Software for Collecting Sign Language Gestures
Sign(ed) languages use gestures, such as hand or head movements, for communication. Sign language recognition is an assistive technology for individuals with hearing disability and its goal is to improve such individuals…
DiversitySign Language RecognitionGeoDiv: Framework For Measuring Geographical Diversity In Text-To-Image Models
Text-to-image (T2I) models are rapidly gaining popularity, yet their outputs often lack geographical diversity, reinforce stereotypes, and misrepresent regions. Given their broad reach, it is critical to rigorously evalu…