paper-with-me

Papers

Physics-based Scene-level Reasoning for Object Pose Estimation in Clutter

2018-06-25 · Chaitanya Mitash, Abdeslam Boularias, Kostas Bekris

This paper focuses on vision-based pose estimation for multiple rigid objects placed in clutter, especially in cases involving occlusions and objects resting on each other. Progress has been achieved recently in object recognition given advancements in deep learning. Nevertheless, such tools typically require a large amount of training data and significant manual effort to label objects. This limits their applicability in robotics, where solutions must scale to a large number of objects and variety of conditions. Moreover, the combinatorial nature of the scenes that could arise from the placement of multiple objects is hard to capture in the training dataset. Thus, the learned models might not produce the desired level of precision required for tasks, such as robotic manipulation. This work proposes an autonomous process for pose estimation that spans from data generation to scene-level reasoning and self-learning. In particular, the proposed framework first generates a labeled dataset for training a Convolutional Neural Network (CNN) for object detection in clutter. These detections are used to guide a scene-level optimization process, which considers the interactions between the different objects present in the clutter to output pose estimates of high precision. Furthermore, confident estimates are used to label online real images from multiple views and re-train the process in a self-learning pipeline. Experimental results indicate that this process is quickly able to identify in cluttered scenes physically-consistent object poses that are more precise than the ones found by reasoning over individual instances of objects. Furthermore, the quality of pose estimates increases over time given the self-learning process.

📄 PDF Abstract BibTeX arXiv:1806.10457

Code (0)

등록된 구현이 없습니다.

Tasks

object-detectionObject DetectionObject RecognitionPose EstimationSelf-Learning

Similar Papers 제목 키워드 기반

NewtPhys: Do Foundation Models Understand Newtonian Physics?

2026-06-02 · Sebastian Cavada, Soumava Paul, Tuan-Hung Vu, Andrei Bursuc 외 arxiv

Previous work has evaluated physics reasoning in foundation models using synthetic or semi-synthetic scenes and visual question-answering tasks. However, these benchmarks emphasize high-level events and lack the visual f…

Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models

2026-03-30 · Nanxi Li, Xiang Wang, Yuanjie Chen, Haode Zhang 외 arxiv

While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in image and video understanding, their ability to comprehend the physical world has become an increasingly important research focu…

STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics-Physics Dual System

2026-05-15 · Zhen Luo, Yixuan Yang, Xudong Xu, Jinkun Hao 외 arxiv

Generating simulation-ready tabletop scenes from task instructions is an intriguing and promising research direction in the field of Embodied AI. However, existing task-to-scene generation methods rely exclusively on lar…

Spatial ReasoningScene Generation

Picasso: Holistic Scene Reconstruction with Physics-Constrained Sampling

2026-02-08 · Xihang Yu, Rajat Talak, Lorenzo Shaikewitz, Luca Carlone arxiv

In the presence of occlusions and measurement noise, geometrically accurate scene reconstructions -- which fit the sensor data -- can still be physically incorrect. For instance, when estimating the poses and shapes of o…

Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering

2024-06-02 · Xingrui Wang, Wufei Ma, Angtian Wang, Shuo Chen 외

For vision-language models (VLMs), understanding the dynamic properties of objects and their interactions within 3D scenes from video is crucial for effective reasoning. In this work, we introduce a video question answer…

counterfactualCounterfactual ReasoningFuture predictionQuestion Answering+1