paper-with-me

Papers

Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis

2025-01-12 · Minu Kim, Kangwook Jang, Hoirin Kim

This paper examines how linguistic similarity affects cross-lingual phonetic representation in speech processing for low-resource languages, emphasizing effective source language selection. Previous cross-lingual research has used various source languages to enhance performance for the target low-resource language without thorough consideration of selection. Our study stands out by providing an in-depth analysis of language selection, supported by a practical approach to assess phonetic proximity among multiple language families. We investigate how within-family similarity impacts performance in multilingual training, which aids in understanding language dynamics. We also evaluate the effect of using phonologically similar languages, regardless of family. For the phoneme recognition task, utilizing phonologically similar languages consistently achieves a relative improvement of 55.6% over monolingual training, even surpassing the performance of a large-scale self-supervised learning model. Multilingual training within the same language family demonstrates that higher phonological similarity enhances performance, while lower similarity results in degraded performance compared to monolingual training.

📄 PDF Abstract BibTeX arXiv:2501.06810

Code (0)

등록된 구현이 없습니다.

Tasks

Phoneme RecognitionSelf-Supervised Learning

Similar Papers 제목 키워드 기반

That Sounds Familiar: an Analysis of Phonetic Representations Transfer Across Languages

2020-05-16 · Piotr Żelasko, Laureano Moro-Velázquez, Mark Hasegawa-Johnson, Odette Scharenborg 외

Only a handful of the world's languages are abundant with the resources that enable practical applications of speech processing technologies. One of the methods to overcome this problem is to use the resources existing i…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Language-universal phonetic encoder for low-resource speech recognition

2023-05-19 · Siyuan Feng, Ming Tu, Rui Xia, Chuanzeng Huang 외

Multilingual training is effective in improving low-resource ASR, which may partially be explained by phonetic representation sharing between languages. In end-to-end (E2E) ASR systems, graphemes are often used as basic …

Decoderspeech-recognitionSpeech Recognition

XLEnt: Mining a Large Cross-lingual Entity Dataset with Lexical-Semantic-Phonetic Word Alignment

2021-04-17 · EMNLP 2021 11 · Ahmed El-Kishky, Adithya Renduchintala, James Cross, Francisco Guzmán 외

Cross-lingual named-entity lexica are an important resource to multilingual NLP tasks such as machine translation and cross-lingual wikification. While knowledge bases contain a large number of entities in high-resource …

Machine TranslationMultilingual NLPTranslationWord Alignment

Acoustic-Phonetic Approach for ASR of Less Resourced Languages Using Monolingual and Cross-Lingual Information

2020-05-01 · LREC 2020 5 · shweta bansal

The exploration of speech processing for endangered languages has substantially increased in the past epoch of time. In this paper, we present the acoustic-phonetic approach for automatic speech recognition (ASR) using m…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Modellingspeech-recognition+1

Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet

2026-06-18 · Milan Miletić, Julie Kallini, Ekaterina Shutova arxiv

Multilingual language models often exhibit performance disparities across languages that can arise as early as the tokenization stage. Widely-used subword tokenization approaches favor high-resource languages, and tokeni…