paper-with-me

Papers

Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining

2026-04-02 · Junxuan Li, Rawal Khirodkar, Chengan He, Zhongshi Jiang, Giljoo Nam, Lingchen Yang, Jihyun Lee, Egor Zakharov, Zhaoen Su, Rinat Abdrashitov, Yuan Dong, Julieta Martinez, Kai Li, Qingyang Tan, Takaaki Shiratori, Matthew Hu, Peihong Guo, Xuhua Huang, Ariyan Zarei, Marco Pesavento, Yichen Xu, He Wen, Teng Deng, Wyatt Borsos, Anjali Thakrar, Jean-Charles Bazin, Carsten Stoll, Ginés Hidalgo, James Booth, Lucy Wang, Xiaowen Ma, Yu Rong, Sairanjith Thalanki, Chen Cao, Christian Häne, Abhishek Kar, Sofien Bouaziz, Jason Saragih, Yaser Sheikh, Shunsuke Saito arxiv

High-quality 3D avatar modeling faces a critical trade-off between fidelity and generalization. On the one hand, multi-view studio data enables high-fidelity modeling of humans with precise control over expressions and poses, but it struggles to generalize to real-world data due to limited scale and the domain gap between the studio environment and the real world. On the other hand, recent large-scale avatar models trained on millions of in-the-wild samples show promise for generalization across a wide range of identities, yet the resulting avatars are often of low-quality due to inherent 3D ambiguities. To address this, we present Large-Scale Codec Avatars (LCA), a high-fidelity, full-body 3D avatar model that generalizes to world-scale populations in a feedforward manner, enabling efficient inference. Inspired by the success of large language models and vision foundation models, we present, for the first time, a pre/post-training paradigm for 3D avatar modeling at scale: we pretrain on 1M in-the-wild videos to learn broad priors over appearance and geometry, then post-train on high-quality curated data to enhance expressivity and fidelity. LCA generalizes across hair styles, clothing, and demographics while providing precise, fine-grained facial expressions and finger-level articulation control, with strong identity preservation. Notably, we observe emergent generalization to relightability and loose garment support to unconstrained inputs, and zero-shot robustness to stylized imagery, despite the absence of direct supervision.

📄 PDF Abstract BibTeX arXiv:2604.02320

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Expressive Telepresence via Modular Codec Avatars

2020-08-26 · ECCV 2020 8 · Hang Chu, Shugao Ma, Fernando de la Torre, Sanja Fidler 외

VR telepresence consists of interacting with another human in a virtual space represented by an avatar. Today most avatars are cartoon-like, but soon the technology will allow video-realistic ones. This paper aims in thi…

Gaussian Pixel Codec Avatars: A Hybrid Representation for Efficient Rendering

2025-12-17 · Divam Gupta, Anuj Pahuja, Nemanja Bartolovic, Tomas Simon 외 arxiv

We present Gaussian Pixel Codec Avatars (GPiCA), photorealistic head avatars that can be generated from multi-view images and efficiently rendered on mobile devices. GPiCA utilizes a unique hybrid representation that com…

Audio- and Gaze-driven Facial Animation of Codec Avatars

2020-08-11 · Alexander Richard, Colin Lea, Shugao Ma, Juergen Gall 외

Codec Avatars are a recent class of learned, photorealistic face models that accurately represent the geometry and texture of a person in 3D (i.e., for virtual reality), and are almost indistinguishable from video. In th…

FiCA: Feed-forward instant Gaussian Codec Avatars from a Single Portrait Image

2026-06-23 · Kim Youwang, Zhengyu Yang, Liuhao Ge, Yu Rong 외 arxiv

We introduce FiCA, a Feed-forward, instant Gaussian Codec Avatar generation pipeline that creates lifelike avatars from a single portrait image. Generating a photorealistic and drivable avatar from just a single image is…

Pixel Codec Avatars

2021-04-09 · CVPR 2021 1 · Shugao Ma, Tomas Simon, Jason Saragih, Dawei Wang 외

Telecommunication with photorealistic avatars in virtual or augmented reality is a promising path for achieving authentic face-to-face communication in 3D over remote physical distances. In this work, we present the Pixe…

Decoder