paper-with-me

Papers

AnySlot: Goal-Conditioned Vision-Language-Action Policies for Zero-Shot Slot-Level Placement

2026-04-12 · Zhaofeng Hu, Sifan Zhou, Qinbo Zhang, Rongtao Xu, Qi Su, Jorge Mendez-Mendz, Ci-Jyun Liang arxiv

Vision-Language-Action (VLA) policies have emerged as a versatile paradigm for generalist robotic manipulation. However, precise object placement under compositional language remains challenging for end-to-end VLA policies. Slot-level placement requires reliable slot grounding and centimeter-level geometric precision. To this end, we propose AnySlot, a framework that reduces compositional complexity by introducing an explicit spatial visual goal between language grounding and control. AnySlot converts language into a visual goal by rendering a spatial marker at the intended slot, then executes this goal with a goal-conditioned VLA policy. This hierarchical design decouples high-level slot selection from low-level execution, improving semantic accuracy and spatial robustness. Furthermore, recognizing the lack of benchmarks for such precision-demanding tasks, we introduce SlotBench, a structured simulation benchmark with nine task categories for evaluating spatial reasoning in slot-level placement. Extensive experiments show that AnySlot significantly outperforms flat VLA baselines and modular grounding methods in zero-shot slot-level placement.

📄 PDF Abstract BibTeX arXiv:2604.10432

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

VLESA: Vision-Language Embodied Safety Agent for Human Activity Monitoring

2026-06-02 · Hanjiang Hu, Yiyuan Pan, Jiaxing Li, Xusheng Luo 외 arxiv

As AI systems increasingly assist humans in physical tasks, ensuring safety becomes paramount -- physical actions carry immediate and irreversible consequences that digital errors do not. We introduce the Vision-Language…

HIQL: Offline Goal-Conditioned RL with Latent States as Actions

2023-07-22 · NeurIPS 2023 11 · Seohong Park, Dibya Ghosh, Benjamin Eysenbach, Sergey Levine

Unsupervised pre-training has recently become the bedrock for computer vision and natural language processing. In reinforcement learning (RL), goal-conditioned RL can potentially provide an analogous self-supervised appr…

Reinforcement Learning (RL)Unsupervised Pre-training

Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning

2025-02-11 · Qingyuan Wu, Jianheng Liu, Jianye Hao, Jun Wang 외

State-of-the-art (SOTA) reinforcement learning (RL) methods have enabled vision-language model (VLM) agents to learn from interaction with online environments without human supervision. However, these methods often strug…

Decision Makingreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

TouchWorld: A Predictive and Reactive Tactile Foundation Model for Dexterous Manipulation

2026-07-08 · Jianyi Zhou, Feiyang Hong, Yunhao Li, Yicheng Zhao 외 arxiv

Dexterous manipulation in everyday environments requires both anticipation and reaction: a robot must predict how contact should evolve while rapidly correcting local errors caused by slip, misalignment, unstable graspin…

PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations

2026-04-30 · Yang Zhang, Jiangyuan Zhao, Chenyou Fan, Fangzheng Yan 외 arxiv

Vision-Language-Action (VLA) models advance robotic control via strong visual-linguistic priors. However, existing VLAs predominantly frame pretraining as supervised behavior cloning, overlooking the fundamental nature o…

Reinforcement Learning