paper-with-me

홈 › Papers

From Signal to Turn: Interactional Friction in Modular Speech-to-Speech Pipelines

2025-12-12 · Tittaya Mairittha, Tanakon Sawanglok, Panuwit Raden, Jirapast Buntub, Thanapat Warunee, Napat Asawachaisuvikrom, Thanaphum Saiwongin arxiv

While voice-based AI systems have achieved remarkable generative capabilities, their interactions often feel conversationally broken. This paper examines the interactional friction that emerges in modular Speech-to-Speech Retrieval-Augmented Generation (S2S-RAG) pipelines. By analyzing a representative production system, we move beyond simple latency metrics to identify three recurring patterns of conversational breakdown: (1) Temporal Misalignment, where system delays violate user expectations of conversational rhythm; (2) Expressive Flattening, where the loss of paralinguistic cues leads to literal, inappropriate responses; and (3) Repair Rigidity, where architectural gating prevents users from correcting errors in real-time. Through system-level analysis, we demonstrate that these friction points should not be understood as defects or failures, but as structural consequences of a modular design that prioritizes control over fluidity. We conclude that building natural spoken AI is an infrastructure design challenge, requiring a shift from optimizing isolated components to carefully choreographing the seams between them.

📄 PDF Abstract BibTeX arXiv:2512.11724

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CSTrader: A Testbed for Language-Grounded Trading in a Community-Driven Virtual Asset Market

2026-06-30 · Yao Shi, Kingfung Luo, Nan Tang, Yuyu Luo arxiv

Niche asset markets, such as Counter-Strike 2 (CS2) weapon skins, are small, volatile, and heavily driven by community discussions and platform rules. These properties make them hard for traditional quantitative models, …

SpeechNet: A Universal Modularized Model for Speech Processing Tasks

2021-05-07 · Yi-Chen Chen, Po-Han Chi, Shu-wen Yang, Kai-Wei Chang 외

There is a wide variety of speech processing tasks ranging from extracting content information from speech signals to generating speech signals. For different tasks, model networks are usually designed and tuned separate…

Multi-Task Learning

Acoustic and Facial Markers of Perceived Conversational Success in Spontaneous Speech

2026-03-03 · Thanushi Withanage, Elizabeth Redcay, Carol Espy-Wilson arxiv

Individuals often align their speaking patterns with their interlocutors, a phenomenon linked to engagement and rapport. While well documented in task-oriented dialogues, less is known about entrainment in naturalistic, …

Current Challenges in Spoken Dialogue Systems and Why They Are Critical for Those Living with Dementia

2019-09-14 · Angus Addlesee, Arash Eshghi, Ioannis Konstas

Dialogue technologies such as Amazon's Alexa have the potential to transform the healthcare industry. However, current systems are not yet naturally interactive: they are often turn-based, have naive end-of-turn detectio…

Diagnosticspeech-recognitionSpeech RecognitionSpoken Dialogue Systems

TELEVAL: A Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios

2025-07-24 · Zehan Li, Hongjie Chen, Qing Wang, Yuxin Zhang 외 arxiv

Spoken Language Models (SLMs) are expected to support natural spoken interaction beyond task completion. However, existing SLM benchmarks primarily evaluate semantic correctness in structured settings and provide limited…