paper-with-me

Papers

GraSP-VLA: Graph-based Symbolic Action Representation for Long-Horizon Planning with VLA Policies

2025-11-06 · Maëlic Neau, Zoe Falomir, Paulo E. Santos, Anne-Gwenn Bosser, Cédric Buche arxiv

Deploying autonomous robots that can learn new skills from demonstrations is an important challenge of modern robotics. Existing solutions often apply end-to-end imitation learning with Vision-Language Action (VLA) models or symbolic approaches with Action Model Learning (AML). On the one hand, current VLA models are limited by the lack of high-level symbolic planning, which hinders their abilities in long-horizon tasks. On the other hand, symbolic approaches in AML lack generalization and scalability perspectives. In this paper we present a new neuro-symbolic approach, GraSP-VLA, a framework that uses a Continuous Scene Graph representation to generate a symbolic representation of human demonstrations. This representation is used to generate new planning domains during inference and serves as an orchestrator for low-level VLA policies, scaling up the number of actions that can be reproduced in a row. Our results show that GraSP-VLA is effective for modeling symbolic representations on the task of automatic planning domain generation from observations. In addition, results on real-world experiments show the potential of our Continuous Scene Graph representation to orchestrate low-level VLA policies in long-horizon tasks.

📄 PDF Abstract BibTeX arXiv:2511.04357

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Grasp Type Revisited: A Modern Perspective on a Classical Feature for Vision

2015-06-01 · CVPR 2015 6 · Yezhou Yang, Cornelia Fermuller, Yi Li, Yiannis Aloimonos

The grasp type provides crucial information about human action. However, recognizing the grasp type in unconstrained scenes is challenging because of the large variations in appearance, occlusions and geometric distorti…

Action SegmentationAction UnderstandingGeneral ClassificationVocal Bursts Type Prediction

CRAFT-E: A Neuro-Symbolic Framework for Embodied Affordance Grounding

2025-12-03 · Zhou Chen, Joe Lin, Carson Bulgin, Sathyanarayanan N. Aakur arxiv

Assistive robots operating in unstructured environments must understand not only what objects are, but what they can be used for. This requires grounding language-based action queries to objects that both afford the requ…

SynHLMA:Synthesizing Hand Language Manipulation for Articulated Object with Discrete Human Object Interaction Representation

2025-10-29 · Wang zhi, Yuyan Liu, Liu Liu, Li Zhang 외 arxiv

Generating hand grasps with language instructions is a widely studied topic that benefits from embodied AI and VR/AR applications. While transferring into hand articulatied object interaction (HAOI), the hand grasps synt…

GraphMERT: Efficient and Scalable Distillation of Reliable Knowledge Graphs from Unstructured Data

2025-10-10 · Margarita Belova, Jiaxin Xiao, Shikhar Tuli, Niraj K. Jha arxiv

Researchers have pursued neurosymbolic artificial intelligence (AI) applications for nearly three decades. A marriage of the neural and symbolic components can lead to rapid advancements in AI. Yet, the field has not rea…

Knowledge Graphs

Graph-based Neural Modules to Inspect Attention-based Architectures: A Position Paper

2022-10-13 · Breno W. Carvalho, Artur D'Avilla Garcez, Luis C. Lamb

Encoder-decoder architectures are prominent building blocks of state-of-the-art solutions for tasks across multiple fields where deep learning (DL) or foundation models play a key role. Although there is a growing commun…

DecoderPosition