paper-with-me

홈 › Papers

LIP-RTVE: An Audiovisual Database for Continuous Spanish in the Wild

2023-11-21 · LREC 2022 6 · David Gimeno-Gómez, Carlos-D. Martínez-Hinarejos

Speech is considered as a multi-modal process where hearing and vision are two fundamentals pillars. In fact, several studies have demonstrated that the robustness of Automatic Speech Recognition systems can be improved when audio and visual cues are combined to represent the nature of speech. In addition, Visual Speech Recognition, an open research problem whose purpose is to interpret speech by reading the lips of the speaker, has been a focus of interest in the last decades. Nevertheless, in order to estimate these systems in the currently Deep Learning era, large-scale databases are required. On the other hand, while most of these databases are dedicated to English, other languages lack sufficient resources. Thus, this paper presents a semi-automatically annotated audiovisual database to deal with unconstrained natural Spanish, providing 13 hours of data extracted from Spanish television. Furthermore, baseline results for both speaker-dependent and speaker-independent scenarios are reported using Hidden Markov Models, a traditional paradigm that has been widely used in the field of Speech Technologies.

📄 PDF Abstract BibTeX arXiv:2311.12457

Code (1)

david-gimeno/lip-rtve 공식 구현

Tasks

Automatic Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Analysis of Visual Features for Continuous Lipreading in Spanish

2023-11-21 · David Gimeno-Gómez, Carlos-D. Martínez-Hinarejos

During a conversation, our brain is responsible for combining information obtained from multiple senses in order to improve our ability to understand the message we are perceiving. Different studies have shown the import…

Lipreadingspeech-recognitionSpeech RecognitionVisual Speech Recognition

Speaker-Adapted End-to-End Visual Speech Recognition for Continuous Spanish

2023-11-21 · David Gimeno-Gómez, Carlos-D. Martínez-Hinarejos

Different studies have shown the importance of visual cues throughout the speech perception process. In fact, the development of audiovisual approaches has led to advances in the field of speech technologies. However, al…

speech-recognitionSpeech RecognitionVisual Speech Recognition

Expression, Affect, Action Unit Recognition: Aff-Wild2, Multi-Task Learning and ArcFace

2019-09-25 · Dimitrios Kollias, Stefanos Zafeiriou

Affective computing has been largely limited in terms of available data resources. The need to collect and annotate diverse in-the-wild datasets has become apparent with the rise of deep learning models, as the default a…

Action Unit DetectionArousal EstimationEmotion RecognitionFacial Expression Recognition (FER)+1

Continuous-Time Audiovisual Fusion with Recurrence vs. Attention for In-The-Wild Affect Recognition

2022-03-24 · Vincent Karas, Mani Kumar Tellamekala, Adria Mallol-Ragolta, Michel Valstar 외

In this paper, we present our submission to 3rd Affective Behavior Analysis in-the-wild (ABAW) challenge. Learningcomplex interactions among multimodal sequences is critical to recognise dimensional affect from in-the-wi…

Arousal EstimationEmotion RecognitionMultimodal Emotion Recognition

STAViS: Spatio-Temporal AudioVisual Saliency Network

2020-01-09 · CVPR 2020 6 · Antigoni Tsiami, Petros Koutras, Petros Maragos

We introduce STAViS, a spatio-temporal audiovisual saliency network that combines spatio-temporal visual and auditory information in order to efficiently address the problem of saliency estimation in videos. Our approach…

Saliency Prediction