paper-with-me

Papers

SpeechGen: Unlocking the Generative Power of Speech Language Models with Prompts

2023-06-03 · Haibin Wu, Kai-Wei Chang, Yuan-Kuei Wu, Hung-Yi Lee

Large language models (LLMs) have gained considerable attention for Artificial Intelligence Generated Content (AIGC), particularly with the emergence of ChatGPT. However, the direct adaptation of continuous speech to LLMs that process discrete tokens remains an unsolved challenge, hindering the application of LLMs for speech generation. The advanced speech LMs are in the corner, as that speech signals encapsulate a wealth of information, including speaker and emotion, beyond textual data alone. Prompt tuning has demonstrated notable gains in parameter efficiency and competitive performance on some speech classification tasks. However, the extent to which prompts can effectively elicit generation tasks from speech LMs remains an open question. In this paper, we present pioneering research that explores the application of prompt tuning to stimulate speech LMs for various generation tasks, within a unified framework called SpeechGen, with around 10M trainable parameters. The proposed unified framework holds great promise for efficiency and effectiveness, particularly with the imminent arrival of advanced speech LMs, which will significantly enhance the capabilities of the framework. The code and demos of SpeechGen will be available on the project website: \url{https://ga642381.github.io/SpeechPrompt/speechgen}

📄 PDF Abstract BibTeX arXiv:2306.02207

Code (0)

등록된 구현이 없습니다.

Tasks

Open-Ended Question Answering

Similar Papers 제목 키워드 기반

MF-Speech: Achieving Fine-Grained and Compositional Control in Speech Generation via Factor Disentanglement

2025-11-15 · Xinyue Yu, Youqing Fang, Pingyu Wu, Guoyang Ye 외 arxiv

Generating expressive and controllable human speech is one of the core goals of generative artificial intelligence, but its progress has long been constrained by two fundamental challenges: the deep entanglement of speec…

Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning

2026-07-02 · Congrui Du, Yang Zhang, Kaizhi Qian, Shiyu Chang arxiv

Instruction tuning for speech language models (SLMs) is substantially more challenging than for text-based large language models (LLMs), as it requires learning a new modality and a wide range of speech-specific instruct…

Hearing Between the Lines: Unlocking the Reasoning Power of LLMs for Speech Evaluation

2026-01-20 · Arjun Chandra, Kevin Miller, Venkatesh Ravichandran, Constantinos Papayiannis 외 arxiv

Large Language Model (LLM) judges exhibit strong reasoning capabilities but are limited to textual content. This leaves current automatic Speech-to-Speech (S2S) evaluation methods reliant on opaque and expensive Audio La…

A roadmap for generative mapping: unlocking the power of generative AI for map-making

2024-10-21 · Sidi Wu, Katharina Henggeler, Yizi Chen, Lorenz Hurni

Maps are broadly relevant across various fields, serving as valuable tools for presenting spatial phenomena and communicating spatial knowledge. However, map-making is still largely confined to those with expertise in GI…

From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping

2025-12-31 · Xu He, Haoxian Zhang, Hejia Chen, Changyuan Zheng 외 arxiv

Audio-driven visual dubbing aims to synchronize a video's lip movements with new speech but is fundamentally challenged by the lack of ideal training data: paired videos differing only in lip motion. Existing methods cir…