paper-with-me

Papers

MAVD: The First Open Large-Scale Mandarin Audio-Visual Dataset with Depth Information

2023-06-04 · Jianrong Wang, Yuchen Huo, Li Liu, Tianyi Xu, Qi Li, Sen Li

Audio-visual speech recognition (AVSR) gains increasing attention from researchers as an important part of human-computer interaction. However, the existing available Mandarin audio-visual datasets are limited and lack the depth information. To address this issue, this work establishes the MAVD, a new large-scale Mandarin multimodal corpus comprising 12,484 utterances spoken by 64 native Chinese speakers. To ensure the dataset covers diverse real-world scenarios, a pipeline for cleaning and filtering the raw text material has been developed to create a well-balanced reading material. In particular, the latest data acquisition device of Microsoft, Azure Kinect is used to capture depth information in addition to the traditional audio signals and RGB images during data acquisition. We also provide a baseline experiment, which could be used to evaluate the effectiveness of the dataset. The dataset and code will be released at https://github.com/SpringHuo/MAVD.

📄 PDF Abstract BibTeX arXiv:2306.02263

Code (1)

springhuo/mavd 공식 구현

Tasks

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

Similar Papers 제목 키워드 기반

End-to-End Mandarin Tone Classification with Short Term Context Information

2021-04-12 · Jiyang Tang, Ming Li

In this paper, we propose an end-to-end Mandarin tone classification method from continuous speech utterances utilizing both the spectrogram and the short-term context information as the input. Both spectrograms and cont…

General Classification

AISHELL-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline

2017-09-16 · Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu 외

An open-source Mandarin speech corpus called AISHELL-1 is released. It is by far the largest corpus which is suitable for conducting the speech recognition research and building speech recognition systems for Mandarin. T…

speech-recognitionSpeech Recognition

AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale

2018-08-31 · Jiayu Du, Xingyu Na, Xuechen Liu, Hui Bu

AISHELL-1 is by far the largest open-source speech corpus available for Mandarin speech recognition research. It was released with a baseline system containing solid training and testing pipelines for Mandarin ASR. In AI…

Chinese Word Segmentationspeech-recognitionSpeech RecognitionTransfer Learning

AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines

2025-09-28 · Cancan Li, Fei Su, Juan Liu, Hui Bu 외 arxiv

Whisper speech recognition is crucial not only for ensuring privacy in sensitive communications but also for providing a critical communication bridge for patients under vocal restraint and enabling discrete interaction …

Audio-Visual Speech Recognition

FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration

2025-01-24 · Kai-Tuo Xu, Feng-Long Xie, Xu Tang, Yao Hu

We present FireRedASR, a family of large-scale automatic speech recognition (ASR) models for Mandarin, designed to meet diverse requirements in superior performance and optimal efficiency across various applications. Fir…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Computational EfficiencyDecoder+3