paper-with-me

Papers

Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs

2026-06-04 · Gio Paik, Hyunseo Shin, Soungmin Lee arxiv

Automatic Speech Recognition (ASR) has become a key technology for human--AI interaction. However, code-switching ASR (CS-ASR) remains particularly challenging due to the severe scarcity of multilingual CS speech resources across diverse language pairs. Existing approaches primarily improve CS-ASR performance through synthetic CS speech generation or pair-specific fine-tuning on limited bilingual datasets. Nevertheless, these approaches face an inherent scalability limitation, as support for CS must be developed separately for language pairs whose number grows combinatorially with the number of supported languages. In this work, we investigate whether CS capabilities learned from a limited set of seen language pairs can generalize to unseen language pairs through model merging and domain generalization methods. Our experiments show that merged bilingual CS-ASR models modestly generalize to unseen language pairs, suggesting limited transfer of bilingual CS capabilities across language pairs.

📄 PDF Abstract BibTeX arXiv:2606.05846

Code (0)

등록된 구현이 없습니다.

Tasks

Domain GeneralizationSpeech Recognition

Similar Papers 제목 키워드 기반

Phonological Features for 0-shot Multilingual Speech Synthesis

2020-08-06 · Marlene Staib, Tian Huey Teh, Alexandra Torresquintero, Devang S Ram Mohan 외

Code-switching---the intra-utterance use of multiple languages---is prevalent across the world. Within text-to-speech (TTS), multilingual models have been found to enable code-switching. By modifying the linguistic input…

Speech Synthesistext-to-speechText to Speech

Learning Multilingual Meta-Embeddings for Code-Switching Named Entity Recognition

2019-08-01 · WS 2019 8 · Genta Indra Winata, Zhaojiang Lin, Pascale Fung

In this paper, we propose Multilingual Meta-Embeddings (MME), an effective method to learn multilingual representations by leveraging monolingual pre-trained embeddings. MME learns to utilize information from these embed…

Language IdentificationMMEnamed-entity-recognitionNamed Entity Recognition+1

Boosting Zero-shot Cross-lingual Retrieval by Training on Artificially Code-Switched Data

2023-05-09 · Robert Litschko, Ekaterina Artemova, Barbara Plank

Transferring information retrieval (IR) models from a high-resource language (typically English) to other languages in a zero-shot fashion has become a widely adopted approach. In this work, we show that the effectivenes…

Cross-Lingual Word EmbeddingsInformation RetrievalRerankingRetrieval+1

Gradual Code-Switching as Inference-Time Cross-Lingual Representational Alignment for LLMs

2025-10-07 · Haneul Yoo, Jiho Jin, Kyunghyun Cho, Alice Oh arxiv

While large language models (LLMs) have achieved notable progress in multilingual settings, their performance remains uneven across languages as LLMs often rely on English-centric latent representations. In this work, we…

Cross-Lingual Transfer

Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language Models across Modalities

2025-10-08 · Rajvee Sheth, Samridhi Raj Sinha, Mahavir Patil, Himanshu Beniwal 외 arxiv

Amidst the rapid advances of large language models (LLMs), most LLMs still struggle with mixed-language inputs, limited Codeswitching (CSW) datasets, and evaluation biases, which hinder their deployment in multilingual s…