paper-with-me

홈 › Papers

i-LAVA: Insights on Low Latency Voice-2-Voice Architecture for Agents

2025-09-25 · Anupam Purwar, Aditya Choudhary arxiv

We experiment with a low-latency, end-to-end voice-to-voice communication model to optimize it for real-time conversational applications. By analyzing components essential to voice to voice (V-2-V) system viz. automatic speech recognition (ASR), text-to-speech (TTS), and dialog management, our work analyzes how to reduce processing time while maintaining high-quality interactions to identify the levers for optimizing V-2-V system. Our work identifies that TTS component which generates life-like voice, full of emotions including natural pauses and exclamations has highest impact on Real time factor (RTF). The experimented V-2-V architecture utilizes CSM1b has the capability to understand tone as well as context of conversation by ingesting both audio and text of prior exchanges to generate contextually accurate speech. We explored optimization of Residual Vector Quantization (RVQ) iterations by the TTS decoder which come at a cost of decrease in the quality of voice generated. Our experimental evaluations also demonstrate that for V-2-V implementations based on CSM most important optimizations can be brought by reducing the number of RVQ Iterations along with the codebooks used in Mimi.

📄 PDF Abstract BibTeX arXiv:2509.20971

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

Low-latency Real-time Voice Conversion on CPU

2023-11-01 · Konstantine Sadov, Matthew Hutter, Asara Near

We adapt the architectures of previous audio manipulation and generation neural networks to the task of real-time any-to-one voice conversion. Our resulting model, LLVC ($\textbf{L}$ow-latency $\textbf{L}$ow-resource $\t…

CPUKnowledge DistillationVoice Conversion

Progressive Voice Trigger Detection: Accuracy vs Latency

2020-10-29 · Siddharth Sigtia, John Bridle, Hywel Richards, Pascal Clark 외

We present an architecture for voice trigger detection for virtual assistants. The main idea in this work is to exploit information in words that immediately follow the trigger phrase. We first demonstrate that by includ…

StreamVC: Real-Time Low-Latency Voice Conversion

2024-01-05 · Yang Yang, Yury Kartynnik, Yunpeng Li, Jiuqiang Tang 외

We present StreamVC, a streaming voice conversion solution that preserves the content and prosody of any source speech while matching the voice timbre from any target speech. Unlike previous approaches, StreamVC produces…

Speech SynthesisVoice Conversion

IntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction Abilities

2024-10-09 · Xin Zhang, Xiang Lyu, Zhihao Du, Qian Chen 외

Current methods of building LLMs with voice interaction capabilities rely heavily on explicit text autoregressive generation before or during speech response generation to maintain content quality, which unfortunately br…

Response Generation

End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users

2026-06-19 · Allan Henry, Solange Rossato, Christian Graff, Sylvain Huet 외 arxiv

Voice control offers an intuitive alternative to manual drone piloting, yet most existing systems rely on rigid command vocabularies that fail to handle the spontaneous, disfluent speech of naive users. This paper addres…

Spoken Language UnderstandingSelf-Supervised LearningKnowledge DistillationIntent Recognition