paper-with-me

Papers

Kaggle Competition: Cantonese Audio-Visual Speech Recognition for In-car Commands

2022-07-06 · Wenliang Dai, Samuel Cahyawijaya, Tiezheng Yu, Elham J Barezi, Pascale Fung

With the rise of deep learning and intelligent vehicles, the smart assistant has become an essential in-car component to facilitate driving and provide extra functionalities. In-car smart assistants should be able to process general as well as car-related commands and perform corresponding actions, which eases driving and improves safety. However, in this research field, most datasets are in major languages, such as English and Chinese. There is a huge data scarcity issue for low-resource languages, hindering the development of research and applications for broader communities. Therefore, it is crucial to have more benchmarks to raise awareness and motivate the research in low-resource languages. To mitigate this problem, we collect a new dataset, namely Cantonese In-car Audio-Visual Speech Recognition (CI-AVSR), for in-car speech recognition in the Cantonese language with video and audio data. Together with it, we propose Cantonese Audio-Visual Speech Recognition for In-car Commands as a new challenge for the community to tackle low-resource speech recognition under in-car scenarios.

📄 PDF Abstract BibTeX arXiv:2207.02663

Code (0)

등록된 구현이 없습니다.

Tasks

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

Similar Papers 제목 키워드 기반

CI-AVSR: A Cantonese Audio-Visual Speech Datasetfor In-car Command Recognition

2022-06-01 · LREC 2022 6 · Wenliang Dai, Samuel Cahyawijaya, Tiezheng Yu, Elham J. Barezi 외

With the rise of deep learning and intelligent vehicles, the smart assistant has become an essential in-car component to facilitate driving and provide extra functionalities. In-car smart assistants should be able to pro…

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

CI-AVSR: A Cantonese Audio-Visual Speech Dataset for In-car Command Recognition

2022-01-11 · Wenliang Dai, Samuel Cahyawijaya, Tiezheng Yu, Elham J. Barezi 외

With the rise of deep learning and intelligent vehicle, the smart assistant has become an essential in-car component to facilitate driving and provide extra functionalities. In-car smart assistants should be able to proc…

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

Automatic Speech Recognition Datasets in Cantonese: A Survey and New Dataset

2022-01-07 · LREC 2022 6 · Tiezheng Yu, Rita Frieske, Peng Xu, Samuel Cahyawijaya 외

Automatic speech recognition (ASR) on low resource languages improves the access of linguistic minorities to technological advantages provided by artificial intelligence (AI). In this paper, we address the problem of dat…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Cultural Vocal Bursts Intensity PredictionPhilosophy+2

HK-LegiCoST: Leveraging Non-Verbatim Transcripts for Speech Translation

2023-06-20 · Cihan Xiao, Henry Li Xinyuan, Jinyi Yang, Dongji Gao 외

We introduce HK-LegiCoST, a new three-way parallel corpus of Cantonese-English translations, containing 600+ hours of Cantonese audio, its standard traditional Chinese transcript, and English translation, segmented and a…

Cross-corpusSentencespeech-recognitionSpeech Recognition+1

Truly Multi-modal YouTube-8M Video Classification with Video, Audio, and Text

2017-06-17 · Zhe Wang, Kingsley Kuan, Mathieu Ravaut, Gaurav Manek 외

The YouTube-8M video classification challenge requires teams to classify 0.7 million videos into one or more of 4,716 classes. In this Kaggle competition, we placed in the top 3% out of 650 participants using released vi…

ClassificationGeneral ClassificationVideo Classification