paper-with-me

Papers

AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines

2025-09-28 · Cancan Li, Fei Su, Juan Liu, Hui Bu, Yulong Wan, Hongbin Suo, Ming Li arxiv

Whisper speech recognition is crucial not only for ensuring privacy in sensitive communications but also for providing a critical communication bridge for patients under vocal restraint and enabling discrete interaction in noise-sensitive environments. The development of Chinese mandarin audio-visual whisper speech recognition is hindered by the lack of large-scale datasets. We present AISHELL6-Whisper, a large-scale open-source audio-visual whisper speech dataset, featuring 30 hours each of whisper speech and parallel normal speech, with synchronized frontal facial videos. Moreover, we propose an audio-visual speech recognition (AVSR) baseline based on the Whisper-Flamingo framework, which integrates a parallel training strategy to align embeddings across speech types, and employs a projection layer to adapt to whisper speech's spectral properties. The model achieves a Character Error Rate (CER) of 4.13% for whisper speech and 1.11% for normal speech in the test set of our dataset, and establishes new state-of-the-art results on the wTIMIT benchmark. The dataset and the AVSR baseline codes are open-sourced at https://zutm.github.io/AISHELL6-Whisper.

📄 PDF Abstract BibTeX arXiv:2509.23833

Code (0)

등록된 구현이 없습니다.

Tasks

Audio-Visual Speech Recognition

Similar Papers 제목 키워드 기반

Decoupling recognition and transcription in Mandarin ASR

2021-08-02 · Jiahong Yuan, Xingyu Cai, Dongji Gao, Renjie Zheng 외

Much of the recent literature on automatic speech recognition (ASR) is taking an end-to-end approach. Unlike English where the writing system is closely related to sound, Chinese characters (Hanzi) represent meaning, not…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

AISHELL-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline

2017-09-16 · Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu 외

An open-source Mandarin speech corpus called AISHELL-1 is released. It is by far the largest corpus which is suitable for conducting the speech recognition research and building speech recognition systems for Mandarin. T…

speech-recognitionSpeech Recognition

AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale

2018-08-31 · Jiayu Du, Xingyu Na, Xuechen Liu, Hui Bu

AISHELL-1 is by far the largest open-source speech corpus available for Mandarin speech recognition research. It was released with a baseline system containing solid training and testing pipelines for Mandarin ASR. In AI…

Chinese Word Segmentationspeech-recognitionSpeech RecognitionTransfer Learning

OCR-Enhanced Multimodal ASR Can Read While Listening

2026-01-26 · Junli Chen, Changli Tang, Yixuan Li, Guangzhi Sun 외 arxiv

Visual information, such as subtitles in a movie, often helps automatic speech recognition. In this paper, we propose Donut-Whisper, an audio-visual ASR model with dual encoder to leverage visual information to improve s…

Audio-Visual Speech RecognitionKnowledge Distillation

Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System

2024-07-13 · Lingwei Meng, Jiawen Kang, Yuejiao Wang, Zengrui Jin 외

Multi-talker speech recognition and target-talker speech recognition, both involve transcription in multi-talker contexts, remain significant challenges. However, existing methods rarely attempt to simultaneously address…

Decoderspeech-recognitionSpeech Recognition