paper-with-me

Papers

GRASS: Unified Generation Model for Speech-to-Semantic Tasks

2023-09-06 · Aobo Xia, Shuyu Lei, Yushu Yang, Xiang Guo, Hua Chai

This paper explores the instruction fine-tuning technique for speech-to-semantic tasks by introducing a unified end-to-end (E2E) framework that generates target text conditioned on a task-related prompt for audio data. We pre-train the model using large and diverse data, where instruction-speech pairs are constructed via a text-to-speech (TTS) system. Extensive experiments demonstrate that our proposed model achieves state-of-the-art (SOTA) results on many benchmarks covering speech named entity recognition, speech sentiment analysis, speech question answering, and more, after fine-tuning. Furthermore, the proposed model achieves competitive performance in zero-shot and few-shot scenarios. To facilitate future work on instruction fine-tuning for speech-to-semantic tasks, we release our instruction dataset and code.

📄 PDF Abstract BibTeX arXiv:2309.02780

Code (0)

등록된 구현이 없습니다.

Tasks

named-entity-recognitionNamed Entity RecognitionQuestion AnsweringSentiment Analysistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation

2025-10-26 · Canxiang Yan, Chunxiang Jin, Dawei Huang, Haibing Yu 외 arxiv

Existing speech models suffer from competing requirements on token representations by understanding and generation tasks. This discrepancy in representation prevents speech language models from performing instruction-bas…

PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models

2024-06-12 · Runyan Yang, Huibao Yang, Xiqing Zhang, Tiantian Ye 외

Recently, there have been attempts to integrate various speech processing tasks into a unified model. However, few previous works directly demonstrated that joint optimization of diverse tasks in multitask speech models …

Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition+1

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling

2026-05-07 · Guanrou Yang, Tian Tan, Qian Chen, Zhikang Niu 외 arxiv

Integrating speech understanding and generation is a pivotal step toward building unified speech models. However, the different representations required for these two tasks currently pose significant compatibility challe…

Self-Supervised LearningSpeech EnhancementVoice Conversion

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

2026-05-28 · Bohan Li, Shi Lian, Hankun Wang, Yiwei Guo 외 arxiv

Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality waveforms. Existing speech tokenizers, however, often fail to satisfy these…

Speech Synthesis

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

2025-08-01 · Le Wang, Jun Wang, Chunyu Qiang, Feng Deng 외 arxiv

We present AudioGen-Omni - a unified approach based on multimodal diffusion transformers (MMDit), capable of generating high-fidelity audio, speech, and song coherently synchronized with the input video. AudioGen-Omni in…

Audio Generation