paper-with-me

Papers

OC16-CE80: A Chinese-English Mixlingual Database and A Speech Recognition Baseline

2016-09-27 · Dong Wang, Zhiyuan Tang, Difei Tang, Qing Chen

We present the OC16-CE80 Chinese-English mixlingual speech database which was released as a main resource for training, development and test for the Chinese-English mixlingual speech recognition (MixASR-CHEN) challenge on O-COCOSDA 2016. This database consists of 80 hours of speech signals recorded from more than 1,400 speakers, where the utterances are in Chinese but each involves one or several English words. Based on the database and another two free data resources (THCHS30 and the CMU dictionary), a speech recognition (ASR) baseline was constructed with the deep neural network-hidden Markov model (DNN-HMM) hybrid system. We then report the baseline results following the MixASR-CHEN evaluation rules and demonstrate that OC16-CE80 is a reasonable data resource for mixlingual research.

📄 PDF Abstract BibTeX arXiv:1609.08412

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Speaker Independent and Multilingual/Mixlingual Speech-Driven Talking Head Generation Using Phonetic Posteriorgrams

2020-06-20 · Huirong Huang, Zhiyong Wu, Shiyin Kang, Dongyang Dai 외

Generating 3D speech-driven talking head has received more and more attention in recent years. Recent approaches mainly have following limitations: 1) most speaker-independent methods need handcrafted features that are t…

Talking Head Generation

Speech Recognition With No Speech Or With Noisy Speech Beyond English

2019-06-17 · Gautam Krishna, Co Tran, Yan Han, Mason Carnahan 외

In this paper we demonstrate continuous noisy speech recognition using connectionist temporal classification (CTC) model on limited Chinese vocabulary using electroencephalography (EEG) features with no speech signal as …

EEGElectroencephalogram (EEG)General ClassificationNoisy Speech Recognition+2

THCHS-30 : A Free Chinese Speech Corpus

2015-12-07 · Dong Wang, Xuewei Zhang

Speech data is crucially important for speech recognition research. There are quite some speech databases that can be purchased at prices that are reasonable for most research institutes. However, for young people who ju…

speech-recognitionSpeech Recognition

Sentiment recognition of Italian elderly through domain adaptation on cross-corpus speech dataset

2022-11-14 · Francesca Gasparini, Alessandra Grossi

The aim of this work is to define a speech emotion recognition (SER) model able to recognize positive, neutral and negative emotions in natural conversations of Italian elderly people. Several datasets for SER are availa…

Cross-corpusDomain AdaptationEmotion RecognitionSpeech Emotion Recognition

OCR-Enhanced Multimodal ASR Can Read While Listening

2026-01-26 · Junli Chen, Changli Tang, Yixuan Li, Guangzhi Sun 외 arxiv

Visual information, such as subtitles in a movie, often helps automatic speech recognition. In this paper, we propose Donut-Whisper, an audio-visual ASR model with dual encoder to leverage visual information to improve s…

Audio-Visual Speech RecognitionKnowledge Distillation