paper-with-me

홈 › Papers

Towards Interactive Intelligence for Digital Humans

2025-12-15 · Yiyi Cai, Xuangeng Chu, Xiwei Gao, Sitong Gong, Yifei Huang, Caixin Kang, Kunhang Li, Haiyang Liu, Ruicong Liu, Yun Liu, Dianwen Ng, Zixiong Su, Erwin Wu, Yuhan Wu, Dingkun Yan, Tianyu Yan, Chang Zeng, Bo Zheng, You Zhou arxiv

We introduce Interactive Intelligence, a novel paradigm of digital human that is capable of personality-aligned expression, adaptive interaction, and self-evolution. To realize this, we present Mio (Multimodal Interactive Omni-Avatar), an end-to-end framework composed of five specialized modules: Thinker, Talker, Face Animator, Body Animator, and Renderer. This unified architecture integrates cognitive reasoning with real-time multimodal embodiment to enable fluid, consistent interaction. Furthermore, we establish a new benchmark to rigorously evaluate the capabilities of interactive intelligence. Extensive experiments demonstrate that our framework achieves superior performance compared to state-of-the-art methods across all evaluated dimensions. Together, these contributions move digital humans beyond superficial imitation toward intelligent interaction.

📄 PDF Abstract BibTeX arXiv:2512.13674

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CartoonAlive: Towards Expressive Live2D Modeling from Single Portraits

2025-07-23 · Chao He, Jianqiang Ren, Jianjing Xiang, Xiejie Shen arxiv

With the rapid advancement of large foundation models, AIGC, cloud rendering, and real-time motion capture technologies, digital humans are now capable of achieving synchronized facial expressions and body movements, eng…

Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans

2026-07-22 · Wentao Jiang, Youchen Xie, Haidi Fan, Yajing Chen 외 hf

Existing co-speech gesture generation methods are predominantly studied in offline settings, where gestures are synthesized from complete speech segments. However, interactive digital humans in real-world scenarios are r…

DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching

2026-06-30 · Meng Chen, Anya Ji, Tsung-Han Wu, Tobias Maringgele 외 arxiv

Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce DigitalCoach, a multimodal dataset of 72 human expert-novice computer use coaching sess…

Visual Grounding

Hi-Reco: High-Fidelity Real-Time Conversational Digital Humans

2025-11-16 · Hongbin Huang, Junwei Li, Tianxin Xie, Zhuang Li 외 arxiv

High-fidelity digital humans are increasingly used in interactive applications, yet achieving both visual realism and real-time responsiveness remains a major challenge. We present a high-fidelity, real-time conversation…

Dialogue GenerationResponse GenerationSpeech Synthesis

Imitating Interactive Intelligence

2020-12-10 · Josh Abramson, Arun Ahuja, Iain Barr, Arthur Brussee 외

A common vision from science fiction is that robots will one day inhabit our physical spaces, sense the world as we do, assist our physical labours, and communicate with us through natural language. Here we study how to …