paper-with-me

Papers

Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch

2026-04-10 · Gabriele Mario Caddeo, Pasquale Marra, Lorenzo Natale arxiv

We propose a multimodal, physically grounded approach for metric-scale amodal object reconstruction and pose estimation under severe hand occlusion. Unlike prior occlusion-aware 3D generation methods that rely only on vision, we leverage physical interaction signals: proprioception provides the posed hand geometry, and multi-contact touch constrains where the object surface must lie, reducing ambiguity in occluded regions. We represent object structure as a pose-aware, camera-aligned signed distance field (SDF) and learn a compact latent space with a Structure-VAE. In this latent space, we train a conditional flow-matching diffusion model, pretraining on vision-only images and finetuning on occluded manipulation scenes while conditioning on visible RGB evidence, occluder/visibility masks, the hand latent representation, and tactile information. Crucially, we incorporate physics-based objectives and differentiable decoder-guidance during finetuning and inference to reduce hand--object interpenetration and to align the reconstructed surface with contact observations. Because our method produces a metric, physically consistent structure estimate, it integrates naturally into existing two-stage reconstruction pipelines, where a downstream module refines geometry and predicts appearance. Experiments in simulation show that adding proprioception and touch substantially improves completion under occlusion and yields physically plausible reconstructions at correct real-world scale compared to vision-only baselines; we further validate transfer by deploying the model on a real humanoid robot with an end-effector different from those used during training.

📄 PDF Abstract BibTeX arXiv:2604.09100

Code (0)

등록된 구현이 없습니다.

Tasks

Pose Estimation3D Generation

Similar Papers 제목 키워드 기반

PhyCo: Learning Controllable Physical Priors for Generative Motion

2026-04-30 · Sriram Narayanan, Ziyu Jiang, Srinivasa Narasimhan, Manmohan Chandraker arxiv

Modern video diffusion models excel at appearance synthesis but still struggle with physical consistency: objects drift, collisions lack realistic rebound, and material responses seldom match their underlying properties.…

Video Generation

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation

2026-04-21 · Feng Jiang, Yang Chen, Kyle Xu, Yuchen Liu 외 arxiv

Recent advances in large-scale video world models have enabled increasingly realistic future prediction, raising the prospect of using generated videos as scalable supervision for robot learning. However, for embodied ma…

Spatial Reasoning

PhysicsNeRF: Physics-Guided 3D Reconstruction from Sparse Views

2025-05-29 · Mohamed Rayan Barhdadi, Hasan Kurban, Hussein Alnuweiri

PhysicsNeRF is a physically grounded framework for 3D reconstruction from sparse views, extending Neural Radiance Fields with four complementary constraints: depth ranking, RegNeRF-style consistency, sparsity priors, and…

3D ReconstructionNeRF

Genie 4D: Semantic-Prior-Guided 4D Dynamic Scene Reconstruction

2026-04-10 · Yiru Yang, Zhuojie Wu, Nishant Kumar Singh, Max Schulthess arxiv

At the intersection of computer vision and robotic perception, 4D reconstruction of dynamic scenes connects low-level geometric sensing with high-level semantic understanding. We present Genie 4D, a framework that turns …

MRD: Using Physically Based Differentiable Rendering to Probe Vision Models for 3D Scene Understanding

2025-12-13 · Benjamin Beilharz, Thomas S. A. Wallis arxiv

While deep learning methods have achieved impressive success in many vision benchmarks, it remains difficult to understand and explain the representations and decisions of these models. Though vision models are typically…

Scene Understanding