paper-with-me

Papers

Two-stage Training for Chinese Dialect Recognition

2019-08-06 · Zongze Ren, Guofu Yang, Shugong Xu

In this paper, we present a two-stage language identification (LID) system based on a shallow ResNet14 followed by a simple 2-layer recurrent neural network (RNN) architecture, which was used for Xunfei (iFlyTek) Chinese Dialect Recognition Challenge and won the first place among 110 teams. The system trains an acoustic model (AM) firstly with connectionist temporal classification (CTC) to recognize the given phonetic sequence annotation and then train another RNN to classify dialect category by utilizing the intermediate features as inputs from the AM. Compared with a three-stage system we further explore, our results show that the two-stage system can achieve high accuracy for Chinese dialects recognition under both short utterance and long utterance conditions with less training time.

📄 PDF Abstract BibTeX arXiv:1908.02284

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationVocal Bursts Valence Prediction

Methods 이 논문이 사용한 방법론

AM 설명 없음

Similar Papers 제목 키워드 기반

Towards Comprehensive Semantic Speech Embeddings for Chinese Dialects

2026-01-12 · Kalvin Chang, Yiwen Shao, Jiahong Li, Dong Yu arxiv

Despite having hundreds of millions of speakers, Chinese dialects lag behind Mandarin in speech and language technologies. Most varieties are primarily spoken, making dialect-to-Mandarin speech-LLMs (large language model…

Speech Recognition

Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis

2025-05-27 · Tianyi Xu, Hongjie Chen, Wang Qing, Lv Hang 외

Large-scale training corpora have significantly improved the performance of ASR models. Unfortunately, due to the relative scarcity of data, Chinese accents and dialects remain a challenge for most ASR models. Recent adv…

Accented Speech RecognitionSelf-Supervised Learningspeech-recognitionSpeech Recognition

Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation

2024-08-01 · Xinhan Di, Zihao Chen, Yunming Liang, Junjie Zheng 외

Large-scale text-to-speech (TTS) models have made significant progress recently.However, they still fall short in the generation of Chinese dialectal speech. Toaddress this, we propose Bailing-TTS, a family of large-scal…

Representation LearningSpeech Synthesistext-to-speechText to Speech

Dolphin-CN-Dialect: Where Chinese Dialects Matter

2026-05-09 · Yangyang Meng, Huihang Zhong, Guodong Lin, Guanbo Wang 외 arxiv

We present Dolphin-CN-Dialect, a streaming-capable ASR model with a focus on Chinese and dialect-rich scenarios. Compared to the previous version, Dolphin-CN-Dialect introduces substantial improvements in data processing…

Speech-Driven End-to-End Language Discrimination towards Chinese Dialects

2026-06-17 · Fan Xu, Jian Luo, MingWen Wang, GuoDong Zhou arxiv

Language discrimination among similar languages, varieties, and dialects is a challenging natural language processing task. The traditional text-driven focus leads to poor results. In this paper, we explore the effective…

Speech Recognition