paper-with-me

홈 › Papers

3D Foundation Models Enable Simultaneous Geometry and Pose Estimation of Grasped Objects

2024-07-14 · Weiming Zhi, Haozhan Tang, Tianyi Zhang, Matthew Johnson-Roberson

Humans have the remarkable ability to use held objects as tools to interact with their environment. For this to occur, humans internally estimate how hand movements affect the object's movement. We wish to endow robots with this capability. We contribute methodology to jointly estimate the geometry and pose of objects grasped by a robot, from RGB images captured by an external camera. Notably, our method transforms the estimated geometry into the robot's coordinate frame, while not requiring the extrinsic parameters of the external camera to be calibrated. Our approach leverages 3D foundation models, large models pre-trained on huge datasets for 3D vision tasks, to produce initial estimates of the in-hand object. These initial estimations do not have physically correct scales and are in the camera's frame. Then, we formulate, and efficiently solve, a coordinate-alignment problem to recover accurate scales, along with a transformation of the objects to the coordinate frame of the robot. Forward kinematics mappings can subsequently be defined from the manipulator's joint angles to specified points on the object. These mappings enable the estimation of points on the held object at arbitrary configurations, enabling robot motion to be designed with respect to coordinates on the grasped objects. We empirically evaluate our approach on a robot manipulator holding a diverse set of real-world objects.

📄 PDF Abstract BibTeX arXiv:2407.10331

Code (0)

등록된 구현이 없습니다.

Tasks

Pose Estimation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

AirSplat: Alignment and Rating for Robust Feed-Forward 3D Gaussian Splatting

2026-03-26 · Minh-Quan Viet Bui, Jaeho Moon, Munchurl Kim arxiv

While 3D Vision Foundation Models (3DVFMs) have demonstrated remarkable zero-shot capabilities in visual geometry estimation, their direct application to generalizable novel view synthesis (NVS) remains challenging. In t…

Novel View Synthesis

DrivingDepth: Sparse-Prompted Pixel-wise Scale Correction for Driving Depth Estimation

2026-06-30 · Chi Huang, Wenhao Zhang, Hang Yin, YuAn Wang 외 arxiv

Dense depth estimation for autonomous driving faces a geometry-scale conflict: depth foundation models deliver pixel-aligned dense visual geometry without reliable metric scale, while projected LiDAR provides metric anch…

Autonomous DrivingDepth Estimation

Face Anything: 4D Face Reconstruction from Any Image Sequence

2026-04-21 · Umut Kocasari, Simon Giebenhain, Richard Shaw, Matthias Nießner arxiv

Accurate reconstruction and tracking of dynamic human faces from image sequences is challenging because non-rigid deformations, expression changes, and viewpoint variations occur simultaneously, creating significant ambi…

Dynamic ReconstructionDepth EstimationPoint Tracking

AIM-SLAM: Dense Monocular SLAM via Adaptive and Informative Multi-View Keyframe Prioritization with Foundation Model

2026-03-05 · Jinwoo Jeon, Dong-Uk Seo, Eungchang Mason Lee, Hyun Myung arxiv

Recent advances in geometric foundation models have emerged as a promising alternative for addressing the challenge of dense reconstruction in monocular visual simultaneous localization and mapping (SLAM). Although geome…

Pose Estimation

GeoSurDepth: Harnessing Foundation Model for Spatial Geometry Consistency-Oriented Self-Supervised Surround-View Depth Estimation

2026-01-09 · Weimin Liu, Wenjun Wang, Joshua H. Meng arxiv

Accurate surround-view depth estimation provides a competitive alternative to laser-based sensors and is essential for 3D scene understanding in autonomous driving. While empirical studies have proposed various approache…

Novel View SynthesisImage ReconstructionScene UnderstandingAutonomous Driving