paper-with-me

Papers

Speech-Driven End-to-End Language Discrimination towards Chinese Dialects

2026-06-17 · Fan Xu, Jian Luo, MingWen Wang, GuoDong Zhou arxiv

Language discrimination among similar languages, varieties, and dialects is a challenging natural language processing task. The traditional text-driven focus leads to poor results. In this paper, we explore the effectiveness of speech-driven features towards language discrimination among Chinese dialects. First, we systematically explore the appropriateness of speech-driven MFCC features towards CNN-based language discrimination. Then, we design an end-to-end speech recognition model based on HMM-DNN to predict Chinese dialect words. We adopt attention to extract the discriminative words related to different Chinese dialects. Finally, through a CNN, we combine the word-level embedding and the MFCC-based features. Evaluation of two benchmark Chinese dialect corpora shows the appropriateness and effectiveness of the proposed speech-driven approach to fine-grained Chinese dialect discrimination compared to the state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2606.18584

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

Low-resource Language Discrimination Towards Chinese Dialects with Transfer learning and Data Augmentation

2026-06-17 · Fan Xu, Yangjie Dan, Keyu Yan, Yong Ma 외 arxiv

Chinese dialects discrimination is a challenging natural language processing task due to scarce annotation resource. In this article, we develop a novel Chinese dialects discrimination framework with transfer learning an…

Speech RecognitionTransfer LearningData Augmentation

Towards Comprehensive Semantic Speech Embeddings for Chinese Dialects

2026-01-12 · Kalvin Chang, Yiwen Shao, Jiahong Li, Dong Yu arxiv

Despite having hundreds of millions of speakers, Chinese dialects lag behind Mandarin in speech and language technologies. Most varieties are primarily spoken, making dialect-to-Mandarin speech-LLMs (large language model…

Speech Recognition

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects

2026-08-08 · Yi Shu, Tianyu Peng, Yingzhuo Deng, Wen Yang 외 arxiv

Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to the scarcity of dialect speech data. Moreover, during dialect adaptat…

Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis

2025-05-27 · Tianyi Xu, Hongjie Chen, Wang Qing, Lv Hang 외

Large-scale training corpora have significantly improved the performance of ASR models. Unfortunately, due to the relative scarcity of data, Chinese accents and dialects remain a challenge for most ASR models. Recent adv…

Accented Speech RecognitionSelf-Supervised Learningspeech-recognitionSpeech Recognition

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

2026-05-31 · Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu 외 arxiv

While End-to-End (E2E) Speech-Large Language Models (Speech-LLMs) are rapidly evolving, their evaluation methodologies remain limited to the era of simple transcription. Existing benchmarks suffer from three critical lim…