paper-with-me

홈 › Papers

H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic Manipulation

2025-11-21 · Yijie Zhu, Rui Shao, Ziyang Liu, Jie He, Jizhihui Liu, Jiuru Wang, Zitong Yu arxiv

Unified video and action prediction models hold great potential for robotic manipulation, as future observations offer contextual cues for planning, while actions reveal how interactions shape the environment. However, most existing approaches treat observation and action generation in a monolithic and goal-agnostic manner, often leading to semantically misaligned predictions and incoherent behaviors. To this end, we propose H-GAR, a Hierarchical interaction framework via Goal-driven observation-Action Refinement.To anchor prediction to the task objective, H-GAR first produces a goal observation and a coarse action sketch that outline a high-level route toward the goal. To enable explicit interaction between observation and action under the guidance of the goal observation for more coherent decision-making, we devise two synergistic modules. (1) Goal-Conditioned Observation Synthesizer (GOS) synthesizes intermediate observations based on the coarse-grained actions and the predicted goal observation. (2) Interaction-Aware Action Refiner (IAAR) refines coarse actions into fine-grained, goal-consistent actions by leveraging feedback from the intermediate observations and a Historical Action Memory Bank that encodes prior actions to ensure temporal consistency. By integrating goal grounding with explicit action-observation interaction in a coarse-to-fine manner, H-GAR enables more accurate manipulation. Extensive experiments on both simulation and real-world robotic manipulation tasks demonstrate that H-GAR achieves state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2511.17079

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HierTOD: A Task-Oriented Dialogue System Driven by Hierarchical Goals

2024-11-11 · Lingbo Mo, Shun Jiang, Akash Maharaj, Bernard Hishamunda 외

Task-Oriented Dialogue (TOD) systems assist users in completing tasks through natural language interactions, often relying on a single-layered workflow structure for slot-filling in public tasks, such as hotel bookings. …

Dialogue ManagementManagementNatural Language UnderstandingResponse Generation+2

A Probabilistic Framework for Hierarchical Goal Recognition

2026-04-24 · Chenyuan Zhang, Katherine Ip, Hamid Rezatofighi, Buser Say 외 arxiv

Goal recognition aims to infer an agent's goal from observations of its behaviour. In realistic settings, recognition can benefit from exploiting hierarchical task structure and reasoning under uncertainty. Planning-base…

ReVoLT: Relational Reasoning and Voronoi Local Graph Planning for Target-driven Navigation

2023-01-06 · Junjia Liu, Jianfei Guo, Zehui Meng, Jingtao Xue

Embodied AI is an inevitable trend that emphasizes the interaction between intelligent entities and the real world, with broad applications in Robotics, especially target-driven navigation. This task requires the robot t…

object-detectionObject DetectionRelational ReasoningRelation Extraction

Inferring Hierarchical Structure in Multi-Room Maze Environments

2023-06-23 · Daria de Tinguy, Toon Van de Maele, Tim Verbelen, Bart Dhoedt

Cognitive maps play a crucial role in facilitating flexible behaviour by representing spatial and conceptual relationships within an environment. The ability to learn and infer the underlying structure of the environment…

Efficient Exploration

Causality-driven Hierarchical Structure Discovery for Reinforcement Learning

2022-10-13 · Shaohui Peng, Xing Hu, Rui Zhang, Ke Tang 외

Hierarchical reinforcement learning (HRL) effectively improves agents' exploration efficiency on tasks with sparse reward, with the guide of high-quality hierarchical structures (e.g., subgoals or options). However, how …

Hierarchical Reinforcement LearningMinecraftreinforcement-learningReinforcement Learning+1