paper-with-me

홈 › Papers

Visual speech recognition: aligning terminologies for better understanding

2017-10-03 · Helen L. Bear, Sarah Taylor

We are at an exciting time for machine lipreading. Traditional research stemmed from the adaptation of audio recognition systems. But now, the computer vision community is also participating. This joining of two previously disparate areas with different perspectives on computer lipreading is creating opportunities for collaborations, but in doing so the literature is experiencing challenges in knowledge sharing due to multiple uses of terms and phrases and the range of methods for scoring results. In particular we highlight three areas with the intention to improve communication between those researching lipreading; the effects of interchanging between speech reading and lipreading; speaker dependence across train, validation, and test splits; and the use of accuracy, correctness, errors, and varying units (phonemes, visemes, words, and sentences) to measure system performance. We make recommendations as to how we can be more consistent.

📄 PDF Abstract BibTeX arXiv:1710.01292

Code (0)

등록된 구현이 없습니다.

Tasks

Lipreadingspeech-recognitionSpeech RecognitionVisual Speech Recognition

Similar Papers 제목 키워드 기반

SlideAVSR: A Dataset of Paper Explanation Videos for Audio-Visual Speech Recognition

2024-01-18 · Hao Wang, Shuhei Kurita, Shuichiro Shimizu, Daisuke Kawahara

Audio-visual speech recognition (AVSR) is a multimodal extension of automatic speech recognition (ASR), using video as a complement to audio. In AVSR, considerable efforts have been directed at datasets for facial featur…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Image Comprehension+3

DiscoverPath: A Knowledge Refinement and Retrieval System for Interdisciplinarity on Biomedical Research

2023-09-04 · Yu-Neng Chuang, Guanchu Wang, Chia-Yuan Chang, Kwei-Herng Lai 외

The exponential growth in scholarly publications necessitates advanced tools for efficient article retrieval, especially in interdisciplinary fields where diverse terminologies are used to describe similar research. Trad…

Articlesnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+5

How to Teach DNNs to Pay Attention to the Visual Modality in Speech Recognition

2020-04-17 · George Sterpu, Christian Saam, Naomi Harte

Audio-Visual Speech Recognition (AVSR) seeks to model, and thereby exploit, the dynamic relationship between a human voice and the corresponding mouth movements. A recently proposed multimodal fusion strategy, AV Align, …

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization

2024-06-18 · Young Jin Ahn, Jungwoo Park, Sangha Park, Jonghyun Choi 외

Visual Speech Recognition (VSR) stands at the intersection of computer vision and speech recognition, aiming to interpret spoken content from visual cues. A prominent challenge in VSR is the presence of homophenes-visual…

Landmark-based LipreadingLipreadingspeech-recognitionSpeech Recognition+1

LaSR: Context-Aware Speech Recognition via Latent Reasoning

2026-05-30 · Heyang Liu, Ziyang Cheng, Jiayi Huang, Wenyang Xiao 외 arxiv

Recent advances in Speech Large Language Models (Speech LLMs) have significantly enhanced spoken language understanding and reasoning. However, their contextual awareness is limited, struggling to perform speech recognit…

Spoken Language UnderstandingSpeech Recognition