paper-with-me

홈 › Papers

Instruction Data Generation and Unsupervised Adaptation for Speech Language Models

2024-06-18 · Vahid Noroozi, Zhehuai Chen, Somshubra Majumdar, Steve Huang, Jagadeesh Balam, Boris Ginsburg

In this paper, we propose three methods for generating synthetic samples to train and evaluate multimodal large language models capable of processing both text and speech inputs. Addressing the scarcity of samples containing both modalities, synthetic data generation emerges as a crucial strategy to enhance the performance of such systems and facilitate the modeling of cross-modal relationships between the speech and text domains. Our process employs large language models to generate textual components and text-to-speech systems to generate speech components. The proposed methods offer a practical and effective means to expand the training dataset for these models. Experimental results show progress in achieving an integrated understanding of text and speech. We also highlight the potential of using unlabeled speech data to generate synthetic samples comparable in quality to those with available transcriptions, enabling the expansion of these models to more languages.

📄 PDF Abstract BibTeX arXiv:2406.12946

Code (0)

등록된 구현이 없습니다.

Tasks

Synthetic Data Generationtext-to-speechText to Speech

Similar Papers 제목 키워드 기반

SLM: Bridge the thin gap between speech and text foundation models

2023-09-30 · Mingqiu Wang, Wei Han, Izhak Shafran, Zelin Wu 외

We present a joint Speech and Language Model (SLM), a multitask, multilingual, and dual-modal model that takes advantage of pretrained foundational speech and language models. SLM freezes the pretrained foundation models…

Instruction FollowingLanguage ModelingLanguage ModellingQuestion Answering+2

VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions

2025-09-09 · Jun Zhan, Mingyang Han, Yuxuan Xie, Chen Wang 외 arxiv

Spoken language models (SLMs) have emerged as a unified paradigm for speech understanding and generation, enabling natural human machine interaction. However, while most progress has focused on semantic accuracy and inst…

Instruction Following

Multimodal speech synthesis architecture for unsupervised speaker adaptation

2018-08-20 · Hieu-Thi Luong, Junichi Yamagishi

This paper proposes a new architecture for speaker adaptation of multi-speaker neural-network speech synthesis systems, in which an unseen speaker's voice can be built using a relatively small amount of speech data witho…

Speech Synthesis

CLARITY: Contextual Linguistic Adaptation and Accent Retrieval for Dual-Bias Mitigation in Text-to-Speech Generation

2025-11-14 · Crystal Min Hui Poon, Pai Chet Ng, Xiaoxiao Miao, Immanuel Jun Kai Loh 외 arxiv

Instruction-guided text-to-speech (TTS) research has reached a maturity level where excellent speech generation quality is possible on demand, yet two coupled biases persist in reducing perceived quality: accent bias, wh…

InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training

2025-03-04 · Dingdong Wang, Jin Xu, Ruihang Chu, Zhifang Guo 외

Recent advancements in speech large language models (SpeechLLMs) have attracted considerable attention. Nonetheless, current methods exhibit suboptimal performance in adhering to speech instructions. Notably, the intelli…

Instruction Followingtext-to-speechText to Speech