paper-with-me

Papers

Comparing heterogeneous visual gestures for measuring the diversity of visual speech signals

2018-05-08 · Helen L. Bear, Richard Harvey

Visual lip gestures observed whilst lipreading have a few working definitions, the most common two are; the visual equivalent of a phoneme' and phonemes which are indistinguishable on the lips'. To date there is no formal definition, in part because to date we have not established a two-way relationship or mapping between visemes and phonemes. Some evidence suggests that visual speech is highly dependent upon the speaker. So here, we use a phoneme-clustering method to form new phoneme-to-viseme maps for both individual and multiple speakers. We test these phoneme to viseme maps to examine how similarly speakers talk visually and we use signed rank tests to measure the distance between individuals. We conclude that broadly speaking, speakers have the same repertoire of mouth gestures, where they differ is in the use of the gestures.

📄 PDF Abstract BibTeX arXiv:1805.02948

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringDiversityLipreading

Similar Papers 제목 키워드 기반

Interpretable Diversity Analysis: Visualizing Feature Representations In Low-Cost Ensembles

2023-02-12 · Tim Whitaker, Darrell Whitley

Diversity is an important consideration in the construction of robust neural network ensembles. A collection of well trained models will generalize better if they are diverse in the patterns they respond to and the predi…

Diversity

VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning

2026-03-01 · Mingkang Dong, Hongyi Cai, Jie Li, Sifan Zhou 외 arxiv

The effectiveness of multimodal instruction tuning depends not only on dataset scale, but critically on whether training samples genuinely require visual reasoning. However, existing instruction datasets often contain a …

Visual Reasoning

EMOTION: Expressive Motion Sequence Generation for Humanoid Robots with In-Context Learning

2024-10-30 · Peide Huang, Yuhan Hu, Nataliya Nechyporenko, Daehwa Kim 외

This paper introduces a framework, called EMOTION, for generating expressive motion sequences in humanoid robots, enhancing their ability to engage in humanlike non-verbal communication. Non-verbal cues such as facial ex…

DiversityIn-Context Learning

SignCol: Open-Source Software for Collecting Sign Language Gestures

2019-10-31 · Mohammad Eslami, Mahdi Karami, Sedigheh Eslami, Solale Tabarestani 외

Sign(ed) languages use gestures, such as hand or head movements, for communication. Sign language recognition is an assistive technology for individuals with hearing disability and its goal is to improve such individuals…

DiversitySign Language Recognition

GeoDiv: Framework For Measuring Geographical Diversity In Text-To-Image Models

2026-02-25 · Abhipsa Basu, Mohana Singh, Shashank Agnihotri, Margret Keuper 외 arxiv

Text-to-image (T2I) models are rapidly gaining popularity, yet their outputs often lack geographical diversity, reinforce stereotypes, and misrepresent regions. Given their broad reach, it is critical to rigorously evalu…