paper-with-me

홈 › Papers

ActionReasoning: Robot Action Reasoning in 3D Space with LLM for Robotic Brick Stacking

2026-02-24 · Guangming Wang, Qizhen Ying, Yixiong Jing, Olaf Wysocki, Brian Sheil arxiv

Classical robotic systems typically rely on custom planners designed for constrained environments. While effective in restricted settings, these systems lack generalization capabilities, limiting the scalability of embodied AI and general-purpose robots. Recent data-driven Vision-Language-Action (VLA) approaches aim to learn policies from large-scale simulation and real-world data. However, the continuous action space of the physical world significantly exceeds the representational capacity of linguistic tokens, making it unclear if scaling data alone can yield general robotic intelligence. To address this gap, we propose ActionReasoning, an LLM-driven framework that performs explicit action reasoning to produce physics-consistent, prior-guided decisions for robotic manipulation. ActionReasoning leverages the physical priors and real-world knowledge already encoded in Large Language Models (LLMs) and structures them within a multi-agent architecture. We instantiate this framework on a tractable case study of brick stacking, where the environment states are assumed to be already accurately measured. The environmental states are then serialized and passed to a multi-agent LLM framework that generates physics-aware action plans. The experiments demonstrate that the proposed multi-agent LLM framework enables stable brick placement while shifting effort from low-level domain-specific coding to high-level tool invocation and prompting, highlighting its potential for broader generalization. This work introduces a promising approach to bridging perception and execution in robotic manipulation by integrating physical reasoning with LLMs.

📄 PDF Abstract BibTeX arXiv:2602.21161

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ActionReasoningBench: Reasoning about Actions with and without Ramification Constraints

2024-06-06 · Divij Handa, Pavel Dolin, Shrinidhi Kumbhar, Tran Cao Son 외

Reasoning about Actions and Change (RAC) has historically played a pivotal role in solving foundational AI problems, such as the frame problem. It has driven advancements in AI fields, such as non-monotonic and commonsen…

DiagnosticHallucinationObject Tracking

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training

2025-12-30 · Yi Liu, Sukai Wang, Dafeng Wei, Xiaowei Cai 외 arxiv

General-purpose robotic systems operating in open-world environments must achieve both broad generalization and high-precision action execution, a combination that remains challenging for existing Vision-Language-Action …

Continuous Control

RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation

2024-06-06 · Jiaming Liu, Mengzhen Liu, Zhenyu Wang, Pengju An 외

A fundamental objective in robot manipulation is to enable models to comprehend visual scenes and execute actions. Although existing Vision-Language-Action (VLA) models for robots can handle a range of basic tasks, they …

Common Sense ReasoningMambaPose PredictionRobot Manipulation+2

LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation

2026-08-04 · Fan Yang, Yuting Su, Xiaobo Wang, Yuncheng You 외 arxiv

World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations and anticipate how a scene will evolve. However, existing WAMs often incur substa…

LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model

2026-01-08 · Zhuoyang Liu, Jiaming Liu, Hao Chen, Jiale Yu 외 arxiv

Vision-Language-Action (VLA) models have recently shown strong generalization, with some approaches seeking to explicitly generate linguistic reasoning traces or predict future observations prior to execution. However, e…