paper-with-me

홈 › Papers

PerceptTwin: Semantic Scene Reconstruction for Iterative LLM Planning and Verification

2026-06-02 · Charlie Gauthier, Sacha Morin, Liam Paull arxiv

Simulation environments are useful for both robot policy learning and planning verification and validation. Traditionally, the process of creating a simulation was onerous. Creating a bespoke simulation environment for each individual environment that a robot would operate in was simply infeasible. In this work, we introduce PerceptTwin, a fully automatic pipeline that constructs interactive simulations directly from semantic scene representations produced by a robot's perception stack. PerceptTwin combines open-vocabulary object maps with 3D asset generation, affordance prediction, and commonsense condition checking. These interactive simulations can be used to validate and refine plans before they are executed on the robot hardware. Borrowing from the AI alignment literature, we also introduce an LLM judge that verifies plan correctness and alignment with human preferences. Experiments show that PerceptTwin feedback allows LLM planners to refine plans, enhance safety, and resist harmful black-box prompting attacks. In our suite of tasks, PerceptTwin improves plan success by an average of approximately 39% for GPT5, GPT5Mini, and GPT5Nano planners. Additionally, PerceptTwin also improves human plan verification by up to 18% on average for plans that fail due to unfilled skill preconditions. Our results demonstrate the potential of open-vocabulary scene simulation from robot perception as a foundation for safer, more reliable robot planning.

📄 PDF Abstract BibTeX arXiv:2606.04226

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning Reconstructability for Drone Aerial Path Planning

2022-09-21 · Yilin Liu, Liqiang Lin, Yue Hu, Ke Xie 외

We introduce the first learning-based reconstructability predictor to improve view and path planning for large-scale 3D urban scene acquisition using unmanned drones. In contrast to previous heuristic approaches, our met…

3D Scene Reconstruction

Multi-robot autonomous 3D reconstruction using Gaussian splatting with Semantic guidance

2024-12-03 · Jing Zeng, Qi Ye, Tianle Liu, Yang Xu 외

Implicit neural representations and 3D Gaussian splatting (3DGS) have shown great potential for scene reconstruction. Recent studies have expanded their applications in autonomous reconstruction through task assignment m…

3DGS3D ReconstructionOpen Vocabulary Semantic SegmentationOpen-Vocabulary Semantic Segmentation+1

Iterative Hybrid Discrete-Continuous Viewpoint Planning for UAV Photogrammetry

2026-08-06 · Alan Grech, Daniel Pisani, Andre Grima, Carl James Debono 외 arxiv

Unmanned aerial vehicle (UAV) photogrammetry requires camera networks that provide sufficient surface coverage, image overlap, parallax, and resolution, yet conventional flight patterns are often poorly adapted to scene …

Panoptic 3D Scene Reconstruction From a Single RGB Image

2021-11-03 · NeurIPS 2021 12 · Manuel Dahnert, Ji Hou, Matthias Nießner, Angela Dai

Understanding 3D scenes from a single image is fundamental to a wide variety of tasks, such as for robotics, motion planning, or augmented reality. Existing works in 3D perception from a single RGB image tend to focus on…

2D Panoptic Segmentation3D Instance Segmentation3D Scene Reconstruction3D Semantic Segmentation+6

NTR: Neural Token Reconstruction for Scene Token Bottleneck in End-to-End Driving

2026-05-29 · Jiahui Li, Jiawei Sun, Zixiang Ren, Ming Liu 외 arxiv

Recent perception-free end-to-end (E2E) autonomous driving methods bypass explicit perception outputs by compressing dense image patch tokens into compact scene tokens for downstream trajectory generation and scoring. Wh…

Representation LearningAutonomous Driving