paper-with-me

홈 › Papers

TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition

2025-09-07 · Tran Nguyen Anh, Truong Dinh Dung, Vo Van Nam, Minh N. H. Nguyen arxiv

Code-switching (CS) presents a significant challenge for general Auto-Speech Recognition (ASR) systems. Existing methods often fail to capture the sub tle phonological shifts inherent in CS scenarios. The challenge is particu larly difficult for language pairs like Vietnamese and English, where both distinct phonological features and the ambiguity arising from similar sound recognition are present. In this paper, we propose a novel architecture for Vietnamese-English CS ASR, a Two-Stage Phoneme-Centric model (TSPC). TSPC adopts a phoneme-centric approach based on an extended Vietnamese phoneme set as an intermediate representation for mixed-lingual modeling, while remaining efficient under low computational-resource constraints. Ex perimental results demonstrate that TSPC consistently outperforms exist ing baselines, including PhoWhisper-base, in Vietnamese-English CS ASR, achieving a significantly lower word error rate of 19.06% with reduced train ing resources. Furthermore, the phonetic-based two-stage architecture en ables phoneme adaptation and language conversion to enhance ASR perfor mance in complex CS Vietnamese-English ASR scenarios.

📄 PDF Abstract BibTeX arXiv:2509.05983

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

VALLR: Visual ASR Language Model for Lip Reading

2025-03-27 · Marshall Thomas, Edward Fish, Richard Bowden

Lip Reading, or Visual Automatic Speech Recognition (V-ASR), is a complex task requiring the interpretation of spoken language exclusively from visual cues, primarily lip movements and facial expressions. This task is es…

Automatic Speech RecognitionLanguage ModelingLanguage ModellingLarge Language Model+3

Phoneme-Level Visual Speech Recognition via Point-Visual Fusion and Language Model Reconstruction

2025-07-25 · Matthew Kit Khinn Teng, Haibo Zhang, Takeshi Saitoh arxiv

Visual Automatic Speech Recognition (V-ASR) is a challenging task that involves interpreting spoken language solely from visual information, such as lip movements and facial expressions. This task is notably challenging …

Visual Speech Recognition

VITS2: Improving Quality and Efficiency of Single-Stage Text-to-Speech with Adversarial Learning and Architecture Design

2023-07-31 · Jungil Kong, Jihoon Park, Beomjeong Kim, Jeongmin Kim 외

Single-stage text-to-speech models have been actively studied recently, and their results have outperformed two-stage pipeline systems. Although the previous single-stage model has made great progress, there is room for …

Computational Efficiencytext-to-speechText to Speech

AraS2P: Arabic Speech-to-Phonemes System

2025-09-27 · Bassam Matar, Mohamed Fayed, Ayman Khalafallah arxiv

This paper describes AraS2P, our speech-to-phonemes system submitted to the Iqra'Eval 2025 Shared Task. We adapted Wav2Vec2-BERT via Two-Stage training strategy. In the first stage, task-adaptive continue pretraining was…

TranUSR: Phoneme-to-word Transcoder Based Unified Speech Representation Learning for Cross-lingual Speech Recognition

2023-05-23 · Hongfei Xue, Qijie Shao, Peikun Chen, Pengcheng Guo 외

UniSpeech has achieved superior performance in cross-lingual automatic speech recognition (ASR) by explicitly aligning latent representations to phoneme units using multi-task self-supervised learning. While the learned …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Representation LearningSelf-Supervised Learning+3