paper-with-me

Papers

Cross-Lingual Interleaving for Speech Language Models

2025-12-01 · Adel Moumen, Guangzhi Sun, Philip C. Woodland arxiv

Spoken Language Models (SLMs) aim to learn linguistic competence directly from speech using discrete units, widening access to Natural Language Processing (NLP) technologies for languages with limited written resources. However, progress has been largely English-centric due to scarce spoken evaluation benchmarks and training data, making cross-lingual learning difficult. We present a cross-lingual interleaving method that mixes speech tokens across languages without textual supervision. We also release an EN-FR training dataset, TinyStories (~42k hours), together with EN-FR spoken StoryCloze and TopicCloze benchmarks for cross-lingual semantic evaluation, both synthetically generated using GPT-4. On 360M and 1B SLMs under matched training-token budgets, interleaving improves monolingual semantic accuracy, enables robust cross-lingual continuation, and strengthens cross-lingual hidden-state alignment. Taken together, these results indicate that cross-lingual interleaving is a simple, scalable route to building multilingual SLMs that understand and converse across languages. All resources will be made open-source to support reproducibility.

📄 PDF Abstract BibTeX arXiv:2512.01865

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving

2024-06-16 · Bhavani Shankar, Preethi Jyothi, Pushpak Bhattacharyya

Code-switching is a widely prevalent linguistic phenomenon in multilingual societies like India. Building speech-to-text models for code-switched speech is challenging due to limited availability of datasets. In this wor…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+3

Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving

2025-05-24 · Jingran Xie, Xiang Li, Hui Wang, Yue Yu 외

Large language models (LLMs) have shown remarkable generalization across tasks, leading to increased interest in integrating speech with LLMs. These speech LLMs (SLLMs) typically use supervised fine-tuning to align speec…

Decoder

CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment

2026-02-23 · Hanwen Liu, Saierdaer Yusuyin, Hao Huang, Zhijian Ou arxiv

Large-language-model (LLM)-based text-to-speech (TTS) systems can generate natural speech, but most are not designed for low-latency dual-streaming synthesis. High-quality dual-streaming TTS depends on accurate text--spe…

Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-training

2025-04-29 · Linjuan Wu, Haoran Wei, Huan Lin, TianHao Li 외

Large language models (LLMs) exhibit remarkable multilingual capabilities despite English-dominated pre-training, attributed to cross-lingual mechanisms during pre-training. Existing methods for enhancing cross-lingual t…

Cross-Lingual TransferData AugmentationSemantic Retrieval

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation

2024-12-28 · Ji-Hoon Kim, Hong-Sun Yang, Yoon-Cheol Ju, Il-Hwan Kim 외

The goal of this work is to generate natural speech in multiple languages while maintaining the same speaker identity, a task known as cross-lingual speech synthesis. A key challenge of cross-lingual speech synthesis is …

Speech Synthesis