paper-with-me

Papers

Enhancing Multilingual Speech Generation and Recognition Abilities in LLMs with Constructed Code-switched Data

2024-09-17 · Jing Xu, Daxin Tan, Jiaqi Wang, Xiao Chen

While large language models (LLMs) have been explored in the speech domain for both generation and recognition tasks, their applications are predominantly confined to the monolingual scenario, with limited exploration in multilingual and code-switched (CS) contexts. Additionally, speech generation and recognition tasks are often handled separately, such as VALL-E and Qwen-Audio. In this paper, we propose a MutltiLingual MultiTask (MLMT) model, integrating multilingual speech generation and recognition tasks within the single LLM. Furthermore, we develop an effective data construction approach that splits and concatenates words from different languages to equip LLMs with CS synthesis ability without relying on CS data. The experimental results demonstrate that our model outperforms other baselines with a comparable data scale. Furthermore, our data construction approach not only equips LLMs with CS speech synthesis capability with comparable speaker consistency and similarity to any given speaker, but also improves the performance of LLMs in multilingual speech generation and recognition tasks.

📄 PDF Abstract BibTeX arXiv:2409.10969

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities

2024-10-11 · Aulia Adila, Dessi Lestari, Ayu Purwarianti, Dipta Tanaya 외

An ideal speech recognition model has the capability to transcribe speech accurately under various characteristics of speech signals, such as speaking style (read and spontaneous), speech context (formal and informal), a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Enhancing Multilingual Speech Recognition through Language Prompt Tuning and Frame-Level Language Adapter

2023-09-18 · Song Li, Yongbin You, Xuezhi Wang, Ke Ding 외

Multilingual intelligent assistants, such as ChatGPT, have recently gained popularity. To further expand the applications of multilingual artificial intelligence assistants and facilitate international communication, it …

parameter-efficient fine-tuningspeech-recognitionSpeech Recognition

Enhancing Low-Resource Language and Instruction Following Capabilities of Audio Language Models

2024-09-17 · Potsawee Manakul, Guangzhi Sun, Warit Sirichotedumrong, Kasima Tharnpipitchai 외

Audio language models process audio inputs using textual prompts for tasks like speech recognition and audio captioning. Although built on multilingual pre-trained components, most are trained primarily on English, limit…

Audio captioningInstruction Followingspeech-recognitionSpeech Recognition

Triple X: A LLM-Based Multilingual Speech Recognition System for the INTERSPEECH2025 MLC-SLM Challenge

2025-07-23 · Miaomiao Gao, Xiaoxiao Xiang, Yiwen Guo arxiv

This paper describes our Triple X speech recognition system submitted to Task 1 of the Multi-Lingual Conversational Speech Language Modeling (MLC-SLM) Challenge. Our work focuses on optimizing speech recognition accuracy…

Speech Recognition

Improving Code-Switching Speech Recognition with TTS Data Augmentation

2026-01-02 · Yue Heng Yeo, Yuchen Hu, Shreyas Gopal, Yizhou Peng 외 arxiv

Automatic speech recognition (ASR) for conversational code-switching speech remains challenging due to the scarcity of realistic, high-quality labeled speech data. This paper explores multilingual text-to-speech (TTS) mo…

Speech RecognitionData Augmentation