Kaggle Competition: Cantonese Audio-Visual Speech Recognition for In-car Commands
With the rise of deep learning and intelligent vehicles, the smart assistant has become an essential in-car component to facilitate driving and provide extra functionalities. In-car smart assistants should be able to process general as well as car-related commands and perform corresponding actions, which eases driving and improves safety. However, in this research field, most datasets are in major languages, such as English and Chinese. There is a huge data scarcity issue for low-resource languages, hindering the development of research and applications for broader communities. Therefore, it is crucial to have more benchmarks to raise awareness and motivate the research in low-resource languages. To mitigate this problem, we collect a new dataset, namely Cantonese In-car Audio-Visual Speech Recognition (CI-AVSR), for in-car speech recognition in the Cantonese language with video and audio data. Together with it, we propose Cantonese Audio-Visual Speech Recognition for In-car Commands as a new challenge for the community to tackle low-resource speech recognition under in-car scenarios.
Code (0)
등록된 구현이 없습니다.
Tasks
Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech RecognitionSimilar Papers 제목 키워드 기반
CI-AVSR: A Cantonese Audio-Visual Speech Datasetfor In-car Command Recognition
With the rise of deep learning and intelligent vehicles, the smart assistant has become an essential in-car component to facilitate driving and provide extra functionalities. In-car smart assistants should be able to pro…
Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech RecognitionCI-AVSR: A Cantonese Audio-Visual Speech Dataset for In-car Command Recognition
With the rise of deep learning and intelligent vehicle, the smart assistant has become an essential in-car component to facilitate driving and provide extra functionalities. In-car smart assistants should be able to proc…
Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech RecognitionAutomatic Speech Recognition Datasets in Cantonese: A Survey and New Dataset
Automatic speech recognition (ASR) on low resource languages improves the access of linguistic minorities to technological advantages provided by artificial intelligence (AI). In this paper, we address the problem of dat…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Cultural Vocal Bursts Intensity PredictionPhilosophy+2HK-LegiCoST: Leveraging Non-Verbatim Transcripts for Speech Translation
We introduce HK-LegiCoST, a new three-way parallel corpus of Cantonese-English translations, containing 600+ hours of Cantonese audio, its standard traditional Chinese transcript, and English translation, segmented and a…
Cross-corpusSentencespeech-recognitionSpeech Recognition+1Truly Multi-modal YouTube-8M Video Classification with Video, Audio, and Text
The YouTube-8M video classification challenge requires teams to classify 0.7 million videos into one or more of 4,716 classes. In this Kaggle competition, we placed in the top 3% out of 650 participants using released vi…
ClassificationGeneral ClassificationVideo Classification