paper-with-me

홈 › Papers

SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models

2025-12-05 · Haowen Liu, Shaoxiong Yao, Haonan Chen, Jiawei Gao, Jiayuan Mao, Jia-Bin Huang, Yilun Du arxiv

Vision-Language Models (VLMs) exhibit remarkable common-sense and semantic reasoning capabilities. However, they lack a grounded understanding of physical dynamics. This limitation arises from training VLMs on static internet-scale visual-language data that contain no causal interactions or action-conditioned changes. Consequently, it remains challenging to leverage VLMs for fine-grained robotic manipulation tasks that require physical understanding, reasoning, and corresponding action planning. To overcome this, we present SIMPACT, a test-time, SIMulation-enabled ACTion Planning framework that equips VLMs with physical reasoning through simulation-in-the-loop world modeling, without requiring any additional training. From a single RGB-D observation, SIMPACT efficiently constructs physics simulations, enabling the VLM to propose informed actions, observe simulated rollouts, and iteratively refine its reasoning. By integrating language reasoning with physics prediction, our simulation-enabled VLM can understand contact dynamics and action outcomes in a physically grounded way. Our method demonstrates state-of-the-art performance on five challenging, real-world rigid-body and deformable manipulation tasks that require fine-grained physical reasoning, outperforming existing general-purpose robotic manipulation models. Our results demonstrate that embedding physics understanding via efficient simulation into VLM reasoning at test time offers a promising path towards generalizable embodied intelligence. Project webpage can be found at https://simpact-bot.github.io

📄 PDF Abstract BibTeX arXiv:2512.05955

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

$\texttt{BluePrint}$: A Social Media User Dataset for LLM Persona Evaluation and Training

2025-09-27 · Aurélien Bück-Kaeffer, Je Qin Chooi, Dan Zhao, Maximilian Puelma Touzel 외 arxiv

Large language models (LLMs) offer promising capabilities for simulating social media dynamics at scale, enabling studies that would be ethically or logistically challenging with human subjects. However, the field lacks …

RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation

2025-06-07 · Songhao Han, Boxiang Qiu, Yue Liao, Siyuan Huang 외

Recent advances in vision-language models (VLMs) have enabled instruction-conditioned robotic systems with improved generalization. However, most existing work focuses on reactive System 1 policies, underutilizing VLMs' …

Vision-Language-Action

HARU: Haptic Augmented Reality-Assisted User-Centric Industrial Network Planning

2022-06-24 · Qi Liao, Tianlun Hu, Nikolaj Marchenko, Peter Kulics 외

To support Industry 4.0 applications with haptics and human-machine interaction, 6G requires a new framework that is fully autonomous, visual, and interactive. In this paper, we provide an end-to-end solution, HARU, for …

Camera RelocalizationSensor Fusion

PINNOCHIO: Physics-Informed Neural Network for Coupled Hyperelastic Interface-Volume Simulation in Orthognathic Surgery

2026-06-01 · Jungwook Lee, Daeseung Kim, Kevin Gu, Zhangfeng Hu 외 arxiv

Predicting patient-specific facial soft-tissue deformation is critical for iterative orthognathic surgery planning. However, current computational methods face a strict accuracy-efficiency trade-off: high-fidelity Finite…

Large Language Model Enhanced Differentiable Trajectory Planning for IoT-Enabled Autonomous Driving

2026-07-11 · Shihao Zhang, Jing Yang, Ziyu Song, Zheng Lin 외 arxiv

Autonomous driving planning is a key component of IoT-enabled intelligent transportation systems, requiring vehicles to generate safe, efficient, and executable trajectories in complex urban environments from multi-sourc…

Trajectory PlanningAutonomous DrivingData Augmentation