paper-with-me

홈 › Papers

GenFusion: Feed-forward Human Performance Capture via Progressive Canonical Space Updates

2026-03-30 · Youngjoong Kwon, Yao He, Heejung Choi, Chen Geng, Zhengmao Liu, Jiajun Wu, Ehsan Adeli arxiv

We present a feed-forward human performance capture method that renders novel views of a performer from a monocular RGB stream. A key challenge in this setting is the lack of sufficient observations, especially for unseen regions. Assuming the subject moves continuously over time, we take advantage of the fact that more body parts become observable by maintaining a canonical space that is progressively updated with each incoming frame. This canonical space accumulates appearance information over time and serves as a context bank when direct observations are missing in the current live frame. To effectively utilize this context while respecting the deformation of the live state, we formulate the rendering process as probabilistic regression. This resolves conflicts between past and current observations, producing sharper reconstructions than deterministic regression approaches. Furthermore, it enables plausible synthesis even in regions with no prior observations. Experiments on in-domain (4D-Dress) and out-of-distribution (MVHumanNet) datasets demonstrate the effectiveness of our approach.

📄 PDF Abstract BibTeX arXiv:2603.28997

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GenFusion: Closing the Loop between Reconstruction and Generation via Videos

2025-03-27 · CVPR 2025 1 · Sibo Wu, Congrong Xu, Binbin Huang, Andreas Geiger 외

Recently, 3D reconstruction and generation have demonstrated impressive novel view synthesis results, achieving high fidelity and efficiency. However, a notable conditioning gap can be observed between these two fields, …

3D Generation3D Reconstruction3D Scene ReconstructionNovel View Synthesis

Recurrence is required to capture the representational dynamics of the human visual system

2019-03-14 · Tim C. Kietzmann, Courtney J Spoerer, Lynn Sörensen, Radoslaw M. Cichy 외

The human visual system is an intricate network of brain regions that enables us to recognize the world around us. Despite its abundant lateral and feedback connections, object processing is commonly viewed and studied a…

Transformer Feed-Forward Layers Are Key-Value Memories

2020-12-29 · EMNLP 2021 11 · Mor Geva, Roei Schuster, Jonathan Berant, Omer Levy

Feed-forward layers constitute two-thirds of a transformer model's parameters, yet their role in the network remains under-explored. We show that feed-forward layers in transformer-based language models operate as key-va…

GRAFT: Geometric Refinement and Fitting Transformer for Human Scene Reconstruction

2026-04-21 · Pradyumna YM, Yuxuan Xue, Yue Chen, Nikita Kister 외 arxiv

Reconstructing physically plausible 3D human-scene interactions (HSI) from a single image currently presents a trade-off: optimization based methods offer accurate contact but are slow (~20s), while feed-forward approach…

HumanRAM: Feed-forward Human Reconstruction and Animation Model using Transformers

2025-06-03 · Zhiyuan Yu, Zhe Li, Hujun Bao, Can Yang 외

3D human reconstruction and animation are long-standing topics in computer graphics and vision. However, existing methods typically rely on sophisticated dense-view capture and/or time-consuming per-subject optimization …

3D Human ReconstructionDecoder