paper-with-me

홈 › Papers

Think, Verbalize, then Speak: Bridging Complex Thoughts and Comprehensible Speech

2025-09-19 · Sang Hoon Woo, Sehun Lee, Kang-wook Kim, Gunhee Kim arxiv

Spoken dialogue systems increasingly employ large language models (LLMs) to leverage their advanced reasoning capabilities. However, direct application of LLMs in spoken communication often yield suboptimal results due to mismatches between optimal textual and verbal delivery. While existing approaches adapt LLMs to produce speech-friendly outputs, their impact on reasoning performance remains underexplored. In this work, we propose Think-Verbalize-Speak, a framework that decouples reasoning from spoken delivery to preserve the full reasoning capacity of LLMs. Central to our method is verbalizing, an intermediate step that translates thoughts into natural, speech-ready text. We also introduce ReVerT, a latency-efficient verbalizer based on incremental and asynchronous summarization. Experiments across multiple benchmarks show that our method enhances speech naturalness and conciseness with minimal impact on reasoning. The project page with the dataset and the source code is available at https://yhytoto12.github.io/TVS-ReVerT

📄 PDF Abstract BibTeX arXiv:2509.16028

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding

2026-05-07 · Zhangquan Chen, Manyuan Zhang, Xinlei Yu, Xiang An 외 arxiv

Dynamic spatial reasoning from monocular video is essential for bridging visual intelligence and the physical world, yet remains challenging for vision-language models (VLMs). Prior approaches either verbalize spatial-te…

Reinforcement LearningSpatial Reasoning

Thinking Before You Speak: A Proactive Test-time Scaling Approach

2025-08-26 · Cong Liu, Wenchang Chai, Hejun Wu, Yan Pan 외 arxiv

Large Language Models (LLMs) often exhibit deficiencies with complex reasoning tasks, such as maths, which we attribute to the discrepancy between human reasoning patterns and those presented in the LLMs' training data. …

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models

2025-10-10 · Donghang Wu, Haoyang Zhang, Jun Chen, Xiangyu 외 arxiv

Real-time Spoken Language Models (SLMs) struggle to leverage Chain-of-Thought (CoT) reasoning due to the prohibitive latency of generating the entire thought process sequentially. Enabling SLMs to think while speaking, s…

Mathematical Reasoning

DISSECT: Diagnosing Where Vision Ends and Language Priors Begin in Scientific VLMs

2026-04-06 · Dikshant Kukreja, Kshitij Sah, Karan Goyal, Mukesh Mohania 외 arxiv

When asked to describe a molecular diagram, a Vision-Language Model correctly identifies ``a benzene ring with an -OH group.'' When asked to reason about the same image, it answers incorrectly. The model can see but it c…

Visual Reasoning

MetricPrompt: Prompting Model as a Relevance Metric for Few-shot Text Classification

2023-06-15 · Hongyuan Dong, Weinan Zhang, Wanxiang Che

Prompting methods have shown impressive performance in a variety of text mining tasks and applications, especially few-shot ones. Despite the promising prospects, the performance of prompting model largely depends on the…

ClassificationFew-Shot Text Classificationtext-classificationText Classification