paper-with-me

Papers

Evaluating Self-supervised Speech Models on a Taiwanese Hokkien Corpus

2023-12-06 · Yi-Hui Chou, Kalvin Chang, Meng-Ju Wu, Winston Ou, Alice Wen-Hsin Bi, Carol Yang, Bryan Y. Chen, Rong-Wei Pai, Po-Yen Yeh, Jo-Peng Chiang, Iu-Tshian Phoann, Winnie Chang, Chenxuan Cui, Noel Chen, Jiatong Shi

Taiwanese Hokkien is declining in use and status due to a language shift towards Mandarin in Taiwan. This is partly why it is a low resource language in NLP and speech research today. To ensure that the state of the art in speech processing does not leave Taiwanese Hokkien behind, we contribute a 1.5-hour dataset of Taiwanese Hokkien to ML-SUPERB's hidden set. Evaluating ML-SUPERB's suite of self-supervised learning (SSL) speech representations on our dataset, we find that model size does not consistently determine performance. In fact, certain smaller models outperform larger ones. Furthermore, linguistic alignment between pretraining data and the target language plays a crucial role.

📄 PDF Abstract BibTeX arXiv:2312.06668

Code (1)

sophia1488/ML-SUPERB-on-TW-HK 공식 구현

Tasks

Self-Supervised Learning

Similar Papers 제목 키워드 기반

Speech-to-Speech Translation For A Real-world Unwritten Language

2022-11-11 · arXiv 2022 10 · Peng-Jen Chen, Kevin Tran, Yilin Yang, Jingfei Du 외

We study speech-to-speech translation (S2ST) that translates speech from one language into another language and focuses on building systems to support languages without standard text writing systems. We use English-Taiwa…

Speech-to-Speech TranslationTranslation

Enhancing Taiwanese Hokkien Dual Translation by Exploring and Standardizing of Four Writing Systems

2024-03-18 · Bo-Han Lu, Yi-Hsuan Lin, En-Shiun Annie Lee, Richard Tzong-Han Tsai

Machine translation focuses mainly on high-resource languages (HRLs), while low-resource languages (LRLs) like Taiwanese Hokkien are relatively under-explored. The study aims to address this gap by developing a dual tran…

Machine TranslationTranslation

CLiFT-ASR: A Cross-Lingual Fine-Tuning Framework for Low-Resource Taiwanese Hokkien Speech Recognition

2025-11-10 · Hung-Yang Sung, Chien-Chun Wang, Kuan-Tang Huang, Tien-Hong Lo 외 arxiv

Automatic speech recognition (ASR) for low-resource languages such as Taiwanese Hokkien is difficult due to the scarcity of annotated data. However, direct fine-tuning on Han-character transcriptions often fails to captu…

Speech Recognition

Breeze Taigi: Benchmarks and Models for Taiwanese Hokkien Speech Recognition and Synthesis

2026-02-26 · Yu-Siang Lan, Chia-Sheng Liu, Yi-Chang Chen, Po-Chun Hsu 외 arxiv

Taiwanese Hokkien (Taigi) presents unique opportunities for advancing speech technology methodologies that can generalize to diverse linguistic contexts. We introduce Breeze Taigi, a comprehensive framework centered on s…

Synthetic Data GenerationSpeech Recognition

TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition

2026-02-25 · Cheng-Yeh Yang, Chien-Chun Wang, Li-Wei Chen, Hung-Shin Lee 외 arxiv

Low-resource automatic speech recognition (ASR) continues to pose significant challenges, primarily due to the limited availability of transcribed data for numerous languages. While a wealth of spoken content is accessib…

Speech Recognition