paper-with-me

Papers

CanVEC - the Canberra Vietnamese-English Code-switching Natural Speech Corpus

2020-05-01 · LREC 2020 5 · Li Nguyen, Christopher Bryant

This paper introduces the Canberra Vietnamese-English Code-switching corpus (CanVEC), an original corpus of natural mixed speech that we semi-automatically annotated with language information, part of speech (POS) tags and Vietnamese translations. The corpus, which was built to inform a sociolinguistic study on language variation and code-switching, consists of 10 hours of recorded speech (87k tokens) between 45 Vietnamese-English bilinguals living in Canberra, Australia. We describe how we collected and annotated the corpus by pipelining several monolingual toolkits to considerably speed up the annotation process. We also describe how we evaluated the automatic annotations to ensure corpus reliability. We make the corpus available for research purposes.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

POS

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

ViMedCSS: A Vietnamese Medical Code-Switching Speech Dataset & Benchmark

2026-02-13 · Tung X. Nguyen, Nhu Vo, Giang-Son Nguyen, Duy Mai Hoang 외 arxiv

Code-switching (CS), which is when Vietnamese speech uses English words like drug names or procedures, is a common phenomenon in Vietnamese medical communication. This creates challenges for Automatic Speech Recognition …

Speech RecognitionDomain Adaptation

TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition

2025-09-07 · Tran Nguyen Anh, Truong Dinh Dung, Vo Van Nam, Minh N. H. Nguyen arxiv

Code-switching (CS) presents a significant challenge for general Auto-Speech Recognition (ASR) systems. Existing methods often fail to capture the sub tle phonological shifts inherent in CS scenarios. The challenge is pa…

Speech Recognition

PhoMT: A High-Quality and Large-Scale Benchmark Dataset for Vietnamese-English Machine Translation

2021-10-23 · EMNLP 2021 11 · Long Doan, Linh The Nguyen, Nguyen Luong Tran, Thai Hoang 외

We introduce a high-quality and large-scale Vietnamese-English parallel dataset of 3.02M sentence pairs, which is 2.9M pairs larger than the benchmark Vietnamese-English machine translation corpus IWSLT15. We conduct exp…

DenoisingMachine TranslationSentenceTranslation

Whisper based Cross-Lingual Phoneme Recognition between Vietnamese and English

2025-08-22 · Nguyen Huu Nhat Minh, Tran Nguyen Anh, Truong Dinh Dung, Vo Van Nam 외 arxiv

Cross-lingual phoneme recognition has emerged as a significant challenge for accurate automatic speech recognition (ASR) when mixing Vietnamese and English pronunciations. Unlike many languages, Vietnamese relies on tona…

Speech Recognition

MTet: Multi-domain Translation for English and Vietnamese

2022-10-11 · Chinh Ngo, Trieu H. Trinh, Long Phan, Hieu Tran 외

We introduce MTet, the largest publicly available parallel corpus for English-Vietnamese translation. MTet consists of 4.2M high-quality training sentence pairs and a multi-domain test set refined by the Vietnamese resea…

Machine TranslationSentenceTranslation