Coupled Iterative Refinement for 6D Multi-Object Pose Estimation
We address the task of 6D multi-object pose: given a set of known 3D objects and an RGB or RGB-D input image, we detect and estimate the 6D pose of each object. We propose a new approach to 6D object pose estimation which consists of an end-to-end differentiable architecture that makes use of geometric knowledge. Our approach iteratively refines both pose and correspondence in a tightly coupled manner, allowing us to dynamically remove outliers to improve accuracy. We use a novel differentiable layer to perform pose refinement by solving an optimization problem we refer to as Bidirectional Depth-Augmented Perspective-N-Point (BD-PnP). Our method achieves state-of-the-art accuracy on standard 6D Object Pose benchmarks. Code is available at https://github.com/princeton-vl/Coupled-Iterative-Refinement.
Code (1)
Tasks
6D Pose Estimation using RGBObjectPose EstimationSimilar Papers 제목 키워드 기반
FlowRefiner: Flow Matching-Based Iterative Refinement for 3D Turbulent Flow Simulation
Accurate autoregressive prediction of 3D turbulent flows remains challenging for neural PDE solvers, as small errors in fine-scale structures can accumulate rapidly over rollout. In this paper, we propose FlowRefiner, a …
Learning to Refine: Spectral-Decoupled Iterative Refinement Framework for Precipitation Nowcasting
Accurate precipitation nowcasting is vital for disaster mitigation, but deep learning methods face a key trade-off: regression models produce over-smoothed, spectrally decaying predictions that blur convective details an…
PRISM: Iterative Cross-Modal Posterior Refinement for Dynamic Text-Attributed Graphs
Dynamic text-attributed graphs (DyTAGs) provide a powerful framework for modeling evolving systems in which node semantics and time-dependent interactions are tightly coupled. Recently, multimodal learning has emerged as…
Representation LearningLink PredictionMulti-Scale Iterative Refinement Network for RGB-D Salient Object Detection
The extensive research leveraging RGB-D information has been exploited in salient object detection. However, salient visual cues appear in various scales and resolutions of RGB images due to semantic gaps at different fe…
Objectobject-detectionObject DetectionRGB-D Salient Object Detection+1Iterative refinement, not training objective, makes HuBERT behave differently from wav2vec 2.0
Self-supervised models for speech representation learning now see widespread use for their versatility and performance on downstream tasks, but the effect of model architecture on the linguistic information learned in th…
Representation Learning