paper-with-me

홈 › Papers

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

2026-04-30 · Junbo Cui, Bokai Xu, Chongyi Wang, Tianyu Yu, Weiyue Sun, Yingjing Xu, Tianran Wang, Zhihui He, Wenshuo Ma, Tianchi Cai, Jiancheng Gui, Luoyuan Zhang, Xian Sun, Fuwei Huang, Moye Chen, Zhuo Lin, Hanyu Liu, Qingxin Gui, Qingzhe Han, Yuyang Wen, Huiping Liu, Rongkang Wang, Yaqi Zhang, Hongliang Wei, Chi Chen, You Li, Kechen Fang, Jie Zhou, Yuxuan Li, Guoyang Zeng, Chaojun Xiao, Yankai Lin, Xu Han, Maosong Sun, Zhiyuan Liu, Yuan Yao arxiv

Recent progress in multimodal large language models (MLLMs) has brought AI capabilities from static offline data processing to real-time streaming interaction, yet they still remain far from human-level multimodal interaction. The key bottlenecks are no longer modality coverage or latency alone, but the interaction paradigm itself. First, perception and response are still separated into alternating phases, preventing models from incorporating new inputs for timely adjustment during generation. Second, most current models remain reactive, responding only to explicit user requests instead of acting proactively in the evolving multimodal environment. We present MiniCPM-o 4.5, our latest effort towards human-like multimodal interaction, which mitigates these gaps by real-time full-duplex omni-modal interaction. It can see, listen, and speak simultaneously in real-time, while also exhibiting proactive behaviors such as issuing reminders or comments based on its continuous understanding of the live scene. The key technique behind MiniCPM-o 4.5 is Omni-Flow, a unified streaming framework that aligns omni-modal inputs and outputs along a shared temporal axis. This formulation converts conventional turn-based interaction into a full-duplex, time-aligned process, enabling simultaneous perception and response and allowing proactive behavior to arise within the same framework. With a total of 9B parameters, MiniCPM-o 4.5 approaches Gemini 2.5 Flash in vision-language capabilities, delivering state-of-the-art open-source performance at its scale. It also surpasses Qwen3-Omni-30B-A3B in omni-modal understanding and delivers better speech generation, with significantly higher computation efficiency. Driven by its efficient architecture design and inference optimization, the model can perform real-time full-duplex omni-modal interaction on edge devices with less than 12GB RAM cost.

📄 PDF Abstract BibTeX arXiv:2604.27393

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

2024-10-23 · Qinglin Zhang, Luyao Cheng, Chong Deng, Qian Chen 외

Full-duplex spoken dialogue systems significantly surpass traditional turn-based dialogue systems, as they allow simultaneous bidirectional communication, closely mirroring human-human interactions. However, achieving lo…

Large Language ModelSpoken Dialogue Systems

Realtime-Venus: A full-duplex interaction system with asynchronous delegation

2026-09-12 · Ruixiang Zhao, Hualei Wang, Renhe Sun, Enzhi Zhou 외 hf

Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and linguistic cues, while video interaction also requires grounding the co…

Question Answering

Metronome: Bound the Cache, Keep the Beat for Real-Time Interaction Model Serving

2026-07-02 · Jiaying Meng, Bojie Li arxiv

Real-time interaction models -- Moshi, MiniCPM-o, Qwen-Omni -- turn serving into a periodic real-time task: on every frame a session ingests streaming audio and must respond by a recurring wall-clock deadline, while its …

EgoMem: Lifelong Memory Agent for Full-duplex Omnimodal Models

2025-09-15 · Yiqun Yao, Naitong Yu, Xiang Li, Xin Jiang 외 arxiv

We introduce EgoMem, the first lifelong memory agent tailored for full-duplex models that process real-time omnimodal streams. EgoMem enables real-time models to recognize multiple users directly from raw audiovisual str…

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction

2026-05-17 · Chaoqun He, Mingyang Xiang, Yingjing Xu, Bokai Xu 외 arxiv

Real-time duplex interaction is essential for multimodal AI systems operating in real-world scenarios, where models must continuously process streaming inputs and respond at appropriate moments. However, most existing mu…