paper-with-me

Papers

RoBoSR: Structured Scene Representations for Embodied Robotic Reasoning

2026-06-23 · Kewei Hu, Wanchan Yu, Fangwen Chen, Jing Jiajian, Zimeng Li, Ying Wei, Tianhao Liu, Michael Zhang, Hanwen Kang arxiv

Despite rapid progress, embodied reasoning under real-world variability remains challenging. Existing approaches rely on demonstration-driven sequential biases, limiting flexibility in open-ended and long-horizon tasks that require structured reasoning over evolving states. We introduce RoBoSR, an intermediate structural representation that formulates manipulation as step-wise state transitions over semantically grounded, object-centric scene graphs. By modeling object states and their spatial relations at the perception-action interface, RoBoSR disentangles high-level task reasoning from raw inputs and enables structured reasoning over preconditions, effects, and goal states. This representation endows the agent with causal reasoning capability, enforcing subtask dependencies and supporting coherent long-horizon task planning. To learn such structure-aware reasoning, we construct Manip-Cognition-1.6M, an open-world dataset that jointly supervises scene understanding, instruction interpretation, and subtask planning across diverse tasks. Across several benchmarks and real-world demonstrations, our method consistently outperforms prompting-based methods and classical TAMP baselines in zero-shot generalization and long-horizon tasks. The results underscore structured intermediate representations as a critical inductive bias for scalable embodied reasoning.

📄 PDF Abstract BibTeX arXiv:2606.24338

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot GeneralizationScene Understanding

Similar Papers 제목 키워드 기반

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

2026-07-13 · Xinghang Li, Jun Guo, Qiwei Li, Long Qian 외 hf

Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coh…

Text-to-Image GenerationScene GenerationVideo GenerationImage Editing

GSR: Learning Structured Reasoning for Embodied Manipulation

2026-02-02 · Kewei Hu, Michael Zhang, Wei Ying, Tianhao Liu 외 arxiv

Despite rapid progress, embodied agents still struggle with long-horizon manipulation that requires maintaining spatial consistency, causal dependencies, and goal constraints. A key limitation of existing approaches is t…

Zero-shot Generalization

TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation

2026-05-07 · Hanyu Zhou, Chuanhao Ma, Gim Hee Lee arxiv

Vision-language-action (VLA) models perform well on training-seen robotic tasks but struggle to generalize to unseen scenes and objects. A key limitation lies in their implicit visual representations, which entangle obje…

GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering

2024-12-19 · Saumya Saxena, Blake Buchanan, Chris Paxton, Bingqing Chen 외

In Embodied Question Answering (EQA), agents must explore and develop a semantic understanding of an unseen environment in order to answer a situated question with confidence. This remains a challenging problem in roboti…

Efficient ExplorationEmbodied Question AnsweringQuestion AnsweringWorld Knowledge

iMaC: Translating Actions into Motion and Contact Images for Embodied World Models

2026-06-08 · Zhenyu Wu, Xiuwei Xu, Yukun Zhou, Yifan Li 외 arxiv

Embodied world models have emerged as a pivotal paradigm for visual robotic decision-making and interactive environment simulation. However, conventional embodied frameworks rely on low-dimensional structured action vect…