paper-with-me

홈 › Papers

Mon3tr: Monocular 3D Telepresence with Pre-built Gaussian Avatars as Amortization

2026-01-12 · Fangyu Lin, Yingdong Hu, Zhening Liu, Yufan Zhuang, Zehong Lin, Jun Zhang arxiv

Immersive telepresence aims to transform human interaction in AR/VR applications by enabling lifelike full-body holographic representations for enhanced remote collaboration. However, existing systems rely on hardware-intensive multi-camera setups and demand high bandwidth for volumetric streaming, limiting their real-time performance on mobile devices. To overcome these challenges, we propose Mon3tr, a novel Monocular 3D telepresence framework that integrates 3D Gaussian splatting (3DGS) based parametric human modeling into telepresence for the first time. Mon3tr adopts an amortized computation strategy, dividing the process into a one-time offline multi-view reconstruction phase to build a user-specific avatar and a monocular online inference phase during live telepresence sessions. A single monocular RGB camera is used to capture body motions and facial expressions in real time to drive the 3DGS-based parametric human model, significantly reducing system complexity and cost. The extracted motion and appearance features are transmitted at < 0.2 Mbps over WebRTC's data channel, allowing robust adaptation to network fluctuations. On the receiver side, e.g., Meta Quest 3, we develop a lightweight 3DGS attribute deformation network to dynamically generate corrective 3DGS attribute adjustments on the pre-built avatar, synthesizing photorealistic motion and appearance at ~ 60 FPS. Extensive experiments demonstrate the state-of-the-art performance of our method, achieving a PSNR of > 28 dB for novel poses, an end-to-end latency of ~ 80 ms, and > 1000x bandwidth reduction compared to point-cloud streaming, while supporting real-time operation from monocular inputs across diverse scenarios. Our demos can be found at https://mon3tr3d.github.io.

📄 PDF Abstract BibTeX arXiv:2601.07518

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ProgressiveAvatars: Progressive Animatable 3D Gaussian Avatars

2026-03-17 · Kaiwen Song, Jinkai Cui, Juyong Zhang arxiv

In practical real-time XR and telepresence applications, network and computing resources fluctuate frequently. Therefore, a progressive 3D representation is needed. To this end, we propose ProgressiveAvatars, a progressi…

TeGA: Texture Space Gaussian Avatars for High-Resolution Dynamic Head Modeling

2025-05-08 · Gengyan Li, Paulo Gotardo, Timo Bolkart, Stephan Garbin 외

Sparse volumetric reconstruction and rendering via 3D Gaussian splatting have recently enabled animatable 3D head avatars that are rendered under arbitrary viewpoints with impressive photorealism. Today, such photoreal a…

4kMotion Estimation

Drivable 3D Gaussian Avatars

2023-11-14 · Wojciech Zielonka, Timur Bagautdinov, Shunsuke Saito, Michael Zollhöfer 외

We present Drivable 3D Gaussian Avatars (D3GA), the first 3D controllable model for human bodies rendered with Gaussian splats. Current photorealistic drivable avatars require either accurate 3D registrations during trai…

3DGS

SpatialAvatar-0: High-Quality 4D Head Avatar with Multi-Stage Reconstruction

2026-06-14 · Yiran Wang, Zeyu Zhang, Yuanming Li, Ziming Wang 외 arxiv

High-quality 4D head avatars from one or a few source portraits are central to telepresence, AR/VR, and digital-human interaction. 3D Gaussian Splatting (3DGS) has emerged as the dominant representation, with two complem…

Expressive Telepresence via Modular Codec Avatars

2020-08-26 · ECCV 2020 8 · Hang Chu, Shugao Ma, Fernando de la Torre, Sanja Fidler 외

VR telepresence consists of interacting with another human in a virtual space represented by an avatar. Today most avatars are cartoon-like, but soon the technology will allow video-realistic ones. This paper aims in thi…