paper-with-me

Papers

A Multi-Purpose Audio-Visual Corpus for Multi-Modal Persian Speech Recognition: the Arman-AV Dataset

2023-01-21 · Javad Peymanfard, Samin Heydarian, Ali Lashini, Hossein Zeinali, Mohammad Reza Mohammadi, Nasser Mozayani

In recent years, significant progress has been made in automatic lip reading. But these methods require large-scale datasets that do not exist for many low-resource languages. In this paper, we have presented a new multipurpose audio-visual dataset for Persian. This dataset consists of almost 220 hours of videos with 1760 corresponding speakers. In addition to lip reading, the dataset is suitable for automatic speech recognition, audio-visual speech recognition, and speaker recognition. Also, it is the first large-scale lip reading dataset in Persian. A baseline method was provided for each mentioned task. In addition, we have proposed a technique to detect visemes (a visual equivalent of a phoneme) in Persian. The visemes obtained by this method increase the accuracy of the lip reading task by 7% relatively compared to the previously proposed visemes, which can be applied to other languages as well.

📄 PDF Abstract BibTeX arXiv:2301.10180

Code (0)

등록된 구현이 없습니다.

Tasks

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Lip ReadingSpeaker Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

Similar Papers 제목 키워드 기반

MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation

2023-03-01 · Mohamed Anwar, Bowen Shi, Vedanuj Goswami, Wei-Ning Hsu 외

We introduce MuAViC, a multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation providing 1200 hours of audio-visual speech in 9 languages. It is fully transcribed and covers 6…

Audio-Visual Speech RecognitionRobust Speech Recognitionspeech-recognitionSpeech Recognition+4

Building a synchronous corpus of acoustic and 3D facial marker data for adaptive audio-visual speech synthesis

2012-05-01 · LREC 2012 5 · Dietmar Schabus, Michael Pucher, Gregor Hofer

We have created a synchronous corpus of acoustic and 3D facial marker data from multiple speakers for adaptive audio-visual text-to-speech synthesis. The corpus contains data from one female and two male speakers and amo…

Audio-Visual Speech RecognitionSpeech RecognitionSpeech Synthesistext-to-speech+3

CLeLfPC: a Large Open Multi-Speaker Corpus of French Cued Speech

2022-06-01 · LREC 2022 6 · Brigitte Bigi, Maryvonne Zimmermann, Carine André

Cued Speech is a communication system developed for deaf people to complement speechreading at the phonetic level with hands. This visual communication mode uses handshapes in different placements near the face in combin…

Transliteration

GameVibe: A Multimodal Affective Game Corpus

2024-06-17 · Matthew Barthet, Maria Kaselimi, Kosmas Pinitas, Konstantinos Makantasis 외

As online video and streaming platforms continue to grow, affective computing research has undergone a shift towards more complex studies involving multiple modalities. However, there is still a lack of readily available…

Diversity

Creating HAVIC: Heterogeneous Audio Visual Internet Collection

2012-05-01 · LREC 2012 5 · Stephanie Strassel, Am Morris, a, Jonathan Fiscus 외

Linguistic Data Consortium and the National Institute of Standards and Technology are collaborating to create a large, heterogeneous annotated multimodal corpus to support research in multimodal event detection and relat…

Event Detection