paper-with-me

홈 › Papers

GOPLA: Generalizable Object Placement Learning via Synthetic Augmentation of Human Arrangement

2025-10-16 · Yao Zhong, Hanzhi Chen, Simon Schaefer, Anran Zhang, Stefan Leutenegger arxiv

Robots are expected to serve as intelligent assistants, helping humans with everyday household organization. A central challenge in this setting is the task of object placement, which requires reasoning about both semantic preferences (e.g., common-sense object relations) and geometric feasibility (e.g., collision avoidance). We present GOPLA, a hierarchical framework that learns generalizable object placement from augmented human demonstrations. A multi-modal large language model translates human instructions and visual inputs into structured plans that specify pairwise object relationships. These plans are then converted into 3D affordance maps with geometric common sense by a spatial mapper, while a diffusion-based planner generates placement poses guided by test-time costs, considering multi-plan distributions and collision avoidance. To overcome data scarcity, we introduce a scalable pipeline that expands human placement demonstrations into diverse synthetic training data. Extensive experiments show that our approach improves placement success rates by 30.04 percentage points over the runner-up, evaluated on positioning accuracy and physical plausibility, demonstrating strong generalization across a wide range of real-world robotic placement scenarios.

📄 PDF Abstract BibTeX arXiv:2510.14627

Code (0)

등록된 구현이 없습니다.

Tasks

Collision Avoidance

Similar Papers 제목 키워드 기반

EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

2023-12-11 · Yi Chen, Yuying Ge, Yixiao Ge, Mingyu Ding 외

The pursuit of artificial general intelligence (AGI) has been accelerated by Multimodal Large Language Models (MLLMs), which exhibit superior reasoning, generalization capabilities, and proficiency in processing multimod…

BenchmarkingHuman-Object Interaction DetectionTask Planning

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

2024-12-05 · Lu Qiu, Yuying Ge, Yi Chen, Yixiao Ge 외

The advent of Multimodal Large Language Models, leveraging the power of Large Language Models, has recently demonstrated superior multimodal understanding and reasoning abilities, heralding a new era for artificial gener…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model

GOPlan: Goal-conditioned Offline Reinforcement Learning by Planning with Learned Models

2023-10-30 · Mianchu Wang, Rui Yang, Xi Chen, Hao Sun 외

Offline Goal-Conditioned RL (GCRL) offers a feasible paradigm for learning general-purpose policies from diverse and multi-task offline datasets. Despite notable recent progress, the predominant offline GCRL methods, mai…

Generative Adversarial Networkreinforcement-learning

MonoPlace3D: Learning 3D-Aware Object Placement for 3D Monocular Detection

2025-04-09 · CVPR 2025 1 · Rishubh Parihar, Srinjay Sarkar, Sarthak Vora, Jogendra Kundu 외

Current monocular 3D detectors are held back by the limited diversity and scale of real-world datasets. While data augmentation certainly helps, it's particularly difficult to generate realistic scene-aware augmented dat…

Data AugmentationDiversitySynthetic Data Generation

EPD: Long-term Memory Extraction, Context-awared Planning and Multi-iteration Decision @ EgoPlan Challenge ICML 2024

2024-07-28 · Letian Shi, Qi Lv, Xiang Deng, Liqiang Nie

In this technical report, we present our solution for the EgoPlan Challenge in ICML 2024. To address the real-world egocentric task planning problem, we introduce a novel planning framework which comprises three stages: …

Decision MakingTask Planning