paper-with-me

홈 › Papers

Automated Speaker Independent Visual Speech Recognition: A Comprehensive Survey

2023-06-14 · Praneeth Nemani, G. Sai Krishna, Supriya Kundrapu

Speaker-independent VSR is a complex task that involves identifying spoken words or phrases from video recordings of a speaker's facial movements. Over the years, there has been a considerable amount of research in the field of VSR involving different algorithms and datasets to evaluate system performance. These efforts have resulted in significant progress in developing effective VSR models, creating new opportunities for further research in this area. This survey provides a detailed examination of the progression of VSR over the past three decades, with a particular emphasis on the transition from speaker-dependent to speaker-independent systems. We also provide a comprehensive overview of the various datasets used in VSR research and the preprocessing techniques employed to achieve speaker independence. The survey covers the works published from 1990 to 2023, thoroughly analyzing each work and comparing them on various parameters. This survey provides an in-depth analysis of speaker-independent VSR systems evolution from 1990 to 2023. It outlines the development of VSR systems over time and highlights the need to develop end-to-end pipelines for speaker-independent VSR. The pictorial representation offers a clear and concise overview of the techniques used in speaker-independent VSR, thereby aiding in the comprehension and analysis of the various methodologies. The survey also highlights the strengths and limitations of each technique and provides insights into developing novel approaches for analyzing visual speech cues. Overall, This comprehensive review provides insights into the current state-of-the-art speaker-independent VSR and highlights potential areas for future research.

📄 PDF Abstract BibTeX arXiv:2306.08314

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech RecognitionSurveyVisual Speech Recognition

Similar Papers 제목 키워드 기반

VieSpeaker: A Large-Scale Vietnamese Speaker Recognition Dataset Beyond Visual Dependency

2026-06-23 · Viet Hoang Pham, Tran Trung Nguyen, Bao Thu Ho, Phuong Tuan Dat 외 arxiv

Speaker recognition has advanced rapidly with large-scale training datasets, yet Vietnamese remains under-resourced, with existing corpora limited in scale and acoustic diversity. Most large-scale datasets rely on facial…

Speaker Recognition

A Novel Speech Feature Fusion Algorithm for Text-Independent Speaker Recognition

2022-12-01 · Biao Ma, Chengben Xu, Ye Zhang

A novel speech feature fusion algorithm with independent vector analysis (IVA) and parallel convolutional neural network (PCNN) is proposed for text-independent speaker recognition. Firstly, some different feature types,…

Speaker RecognitionText-Independent Speaker Recognition

Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition

2025-01-25 · Satwinder Singh, Qianli Wang, Zihan Zhong, Clarion Mendes 외

In this paper, we present a speaker-independent dysarthric speech recognition system, with a focus on evaluating the recently released Speech Accessibility Project (SAP-1005) dataset, which includes speech data from indi…

speech-recognitionSpeech Recognition

LIP-RTVE: An Audiovisual Database for Continuous Spanish in the Wild

2023-11-21 · LREC 2022 6 · David Gimeno-Gómez, Carlos-D. Martínez-Hinarejos

Speech is considered as a multi-modal process where hearing and vision are two fundamentals pillars. In fact, several studies have demonstrated that the robustness of Automatic Speech Recognition systems can be improved …

Automatic Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

Effect of different splitting criteria on the performance of speech emotion recognition

2022-10-26 · Bagus Tris Atmaja, Akira Sasou

Traditional speech emotion recognition (SER) evaluations have been performed merely on a speaker-independent condition; some of them even did not evaluate their result on this condition. This paper highlights the importa…

Emotion RecognitionSentenceSpeech Emotion Recognition