paper-with-me

홈 › Papers

Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech

2024-10-02 · Wonjune Kang, Junteng Jia, Chunyang Wu, Wei Zhou, Egor Lakomkin, Yashesh Gaur, Leda Sari, Suyoun Kim, Ke Li, Jay Mahadeokar, Ozlem Kalinli

This work studies the capabilities of a large language model (LLM) to understand paralinguistic aspects of speech without fine-tuning its weights. We utilize an end-to-end system with a speech encoder, which is trained to produce token embeddings such that the LLM's response to an expressive speech prompt is aligned with its response to a semantically matching text prompt that has also been conditioned on the user's speaking style. This framework enables the encoder to generate tokens that capture both linguistic and paralinguistic information and effectively convey them to the LLM, even when the LLM's weights remain completely frozen. To the best of our knowledge, our work is the first to explore how to induce a frozen LLM to understand more than just linguistic content from speech inputs in a general interaction setting. Experiments demonstrate that our system is able to produce higher quality and more empathetic responses to expressive speech prompts compared to several baselines.

📄 PDF Abstract BibTeX arXiv:2410.01162

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Dual Information Speech Language Models for Emotional Conversations

2025-08-11 · Chun Wang, Chenyang Liu, Wenze Xu, Weihong Deng arxiv

Conversational systems relying on text-based large language models (LLMs) often overlook paralinguistic cues, essential for understanding emotions and intentions. Speech-language models (SLMs), which use speech as input,…

SonicBench: Dissecting the Physical Perception Bottleneck in Large Audio Language Models

2026-01-16 · Yirong Sun, Yanjun Chen, Xin Qiu, Gang Zhang 외 arxiv

Large Audio Language Models (LALMs) excel at semantic and paralinguistic tasks, yet their ability to perceive the fundamental physical attributes of audio such as pitch, loudness, and spatial location remains under-explo…

Relational Reasoning

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation

2025-05-19 · Qiongqiong Wang, Hardik B. Sailor, Tianchi Liu, Ai Ti Aw

Current speech-LLMs exhibit limited capability in contextual reasoning alongside paralinguistic understanding, primarily due to the lack of Question-Answer (QA) datasets that cover both aspects. We propose a novel framew…

Dataset Generation

AzeroS: Extending LLM to Speech with Self-Generated Instruction-Free Tuning

2025-12-31 · Yiwen Shao, Wei Liu, Jiahong Li, Tianzi Wang 외 arxiv

Extending large language models (LLMs) to the speech domain has recently gained significant attention. A typical approach connects a pretrained LLM with an audio encoder through a projection module and trains the resulti…

Resurfacing Paralinguistic Awareness in Large Audio Language Models

2026-03-12 · Hao Yang, Minghan Wang, Tongtong Wu, Lizhen Qu 외 arxiv

Large Audio Language Models (LALMs) have expanded the interaction with human to speech modality, which introduces great interactive potential, due to the paralinguistic cues implicitly indicating the user context. Howeve…