paper-with-me

홈 › Papers

Archon: A Unified Multimodal Model for Holistic Digital Human Generation

2026-05-28 · Chong Bao, Shichen Liu, Lijun Yu, David Futschik, Stylianos Moschoglou, Shefali Srivastava, Ziqian Bai, Feitong Tan, Guofeng Zhang, Zhaopeng Cui, Sean Fanello, Yinda Zhang arxiv

Digital humans are fundamental to immersive interaction, yet creating a unified model for holistic modalities, including text, audio, motion, and visual content, remains an open challenge. In this paper, we present Archon, a fully pretrained, human-centric unified multimodal model for holistic avatar generation. Archon unifies seven modalities with modality-specific tokenizers, and a native autoregressive unified multimodal model pretrained on synchronized modalities and 72 diverse tasks to model holistic joint distributions. To address the token explosion challenge in high-fidelity talking videos, we introduce a memory-efficient semantic video reparameterization, achieving 4x token reduction while preserving fine-grained dynamics, coupled with a semantic-driven video diffusion decoder. We further propose a "Thinking in Modality" that decomposes ambiguous cross-modal tasks into stepwise thinking in an alternative chain of modality, progressively enhancing fidelity and controllability. Extensive experiments demonstrate that Archon achieves superior or comparable performance across diverse digital human generation tasks, validating the effectiveness of our unified framework. Project page: https://zju3dv.github.io/archon/.

📄 PDF Abstract BibTeX arXiv:2605.30311

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Archon: An Architecture Search Framework for Inference-Time Techniques

2024-09-23 · Jon Saad-Falcon, Adrian Gamarra Lafuente, Shlok Natarajan, Nahum Maru 외

Inference-time techniques are emerging as highly effective tools to enhance large language model (LLM) capabilities. However, best practices for developing systems that combine these techniques remain underdeveloped due …

Hyperparameter OptimizationInstruction FollowingLarge Language ModelMath

ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video

2026-07-20 · Xiaozhong Lyu, Gen Li, Zhiyin Qian, Xucong Zhang 외 hf

Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal…

Depth Estimation

MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding

2024-10-29 · YuAn Wang, Di Huang, Yaqi Zhang, Wanli Ouyang 외

Generating lifelike human motions from descriptive texts has experienced remarkable research focus in the recent years, propelled by the emerging requirements of digital humans.Despite impressive advances, existing appro…

DescriptiveLanguage ModelingLanguage ModellingMotion Captioning+2

UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation

2025-05-15 · Yi Li, Haonan Wang, Qixiang Zhang, Boyu Xiao 외

The emergence of unified multimodal understanding and generation models is rapidly attracting attention because of their ability to enhance instruction-following capabilities while minimizing model redundancy. However, t…

DiversityInstruction Following

MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators

2023-06-19 · Yaqi Zhang, Di Huang, Bin Liu, Shixiang Tang 외

Generating realistic human motion from given action descriptions has experienced significant advancements because of the emerging requirement of digital humans. While recent works have achieved impressive results in gene…

Motion Generation