paper-with-me

홈 › Papers

MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models

2026-04-14 · Chung-Ming Chien, Manu Orsini, Eugene Kharitonov, Neil Zeghidour, Karen Livescu, Alexandre Défossez arxiv

Speech-to-speech language models have recently emerged to enhance the naturalness of conversational AI. In particular, full-duplex models are distinguished by their real-time interactivity, including handling of pauses, interruptions, and backchannels. However, improving their factuality remains an open challenge. While scaling the model size could address this gap, it would make real-time inference prohibitively expensive. In this work, we propose MoshiRAG, a modular approach that combines a compact full-duplex interface with selective retrieval to access more powerful knowledge sources. Our asynchronous framework enables the model to identify knowledge-demanding queries and ground its responses in external information. By leveraging the natural temporal gap between response onset and the delivery of core information, the retrieval process can be completed while maintaining a natural conversation flow. With this approach, MoshiRAG achieves factuality comparable to the best publicly released non-duplex speech language models while preserving the interactivity inherent to full-duplex systems. Moreover, our flexible design supports plug-and-play retrieval methods without retraining and demonstrates strong performance on out-of-domain mathematical reasoning tasks.

📄 PDF Abstract BibTeX arXiv:2604.12928

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical Reasoning

Similar Papers 제목 키워드 기반

EgoMem: Lifelong Memory Agent for Full-duplex Omnimodal Models

2025-09-15 · Yiqun Yao, Naitong Yu, Xiang Li, Xin Jiang 외 arxiv

We introduce EgoMem, the first lifelong memory agent tailored for full-duplex models that process real-time omnimodal streams. EgoMem enables real-time models to recognize multiple users directly from raw audiovisual str…

SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation

2024-11-27 · Wenyi Yu, Siyin Wang, Xiaoyu Yang, Xianzhao Chen 외

Full-duplex multimodal large language models (LLMs) provide a unified framework for addressing diverse speech understanding and generation tasks, enabling more natural and seamless human-machine conversations. Unlike tra…

Question AnsweringSpeech Enhancementspeech-recognitionSpeech Recognition+2

What Did I Just Say? Self-Listening for Full-Duplex Speech Models

2026-09-04 · Xuanning Zhou, Junyi Ao, Xiaotong Liu, Tom Ko 외 hf

Full-duplex spoken language models can listen and speak simultaneously, enabling them to handle interruptions and backchannels in human conversation. However, text generation, speech synthesis, and audio playback proceed…

Speech SynthesisText Generation

Realtime-Venus: A full-duplex interaction system with asynchronous delegation

2026-09-12 · Ruixiang Zhao, Hualei Wang, Renhe Sun, Enzhi Zhou 외 hf

Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and linguistic cues, while video interaction also requires grounding the co…

Question Answering

Analysis of Half-Duplex Two-Node Slotted ALOHA Network With Asynchronous Traffic

2023-07-12 · Seyed Ali Hashemian, Farid Ashtiani

Despite the long history of research on slotted ALOHA, the exact analysis of the average delay is still in question as the performance of each node is coupled with the activity of other nodes. In this paper, we consider …