paper-with-me

Papers

Zero-Shot Text-to-Speech for Vietnamese

2025-06-02 · Thi Vu, Linh The Nguyen, Dat Quoc Nguyen

This paper introduces PhoAudiobook, a newly curated dataset comprising 941 hours of high-quality audio for Vietnamese text-to-speech. Using PhoAudiobook, we conduct experiments on three leading zero-shot TTS models: VALL-E, VoiceCraft, and XTTS-V2. Our findings demonstrate that PhoAudiobook consistently enhances model performance across various metrics. Moreover, VALL-E and VoiceCraft exhibit superior performance in synthesizing short sentences, highlighting their robustness in handling diverse linguistic contexts. We publicly release PhoAudiobook to facilitate further research and development in Vietnamese text-to-speech.

📄 PDF Abstract BibTeX arXiv:2506.01322

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

Multilingual LLM Prompting Strategies for Medical English-Vietnamese Machine Translation

2025-09-19 · Nhu Vo, Nu-Uyen-Phuong Le, Dung D. Le, Massimo Piccardi 외 arxiv

Medical English-Vietnamese machine translation (En-Vi MT) is essential for healthcare access and communication in Vietnam, yet Vietnamese remains a low-resource and under-studied language. We systematically evaluate prom…

Machine Translation

VietAIDetector: An Open-Source Zero-Shot Detector for Vietnamese AI-Generated Text

2026-08-26 · Trieu Hai Nguyen, Van-Dung Hoang arxiv

In recent years, distinguishing between AI-generated text and human-written text has remained a challenge. In this paper, we introduce VietAIDetector, an open-source tool designed specifically for detecting Vietnamese AI…

ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport

2026-02-26 · Quoc-Khang Tran, Minh-Thien Nguyen, Nguyen-Khang Pham arxiv

Image-text retrieval has become a fundamental component in intelligent multimedia systems; however, most existing vision-language models are optimized for highresource languages and remain suboptimal for low-resource set…

Cross-Modal RetrievalContrastive LearningText Retrieval

ViSpeechFormer: A Phonemic Approach for Vietnamese Automatic Speech Recognition

2026-02-10 · Khoa Anh Nguyen, Long Minh Hoang, Nghia Hieu Nguyen, Luan Thanh Nguyen 외 arxiv

Vietnamese has a phonetic orthography, where each grapheme corresponds to at most one phoneme and vice versa. Exploiting this high grapheme-phoneme transparency, we propose ViSpeechFormer (\textbf{Vi}etnamese \textbf{Spe…

Speech Recognition

Wanna hear your voice? A sample is all we need!

2024-10-01 · The Hieu Pham, Phuong Thanh Tran Nguyen, Xuan Tho Nguyen, Tan Dat Nguyen 외

Research on audio clue-based target speaker extraction (TSE) has focused on modeling mixtures and reference speech, achieving strong results in English due to abundant datasets. However, cross-lingual properties remain u…

AllSpeech SeparationTarget Speaker Extraction