paper-with-me

Papers

InDex: Empowering VLA Models with Intent-Conditioned Arm-Hand Coordination for Dexterous Manipulation

2026-06-10 · Chuanke Pang, Junyi Huang, Zhijun Zhao, Yaobing Wang, Kun Xu, Xilun Ding arxiv

Pre-trained Vision-Language-Action (VLA) models provide useful semantic and spatial priors, yet their parallel-gripper action interfaces do not specify how those priors should be realized by a dexterous hand. Directly appending finger joints conflates two decisions with different structure: when contact should be established and how a morphology-specific hand trajectory should establish it. We introduce InDex, an intent-conditioned adaptation framework that separates these decisions without discarding full hand supervision. InDex derives a normalized grasp intent from retargeted demonstrations. A first stage predicts synchronized end-effector--intent chunks; conditioned on these predictions, VLA context, and proprioception, a diffusion decoder generates multi-joint hand actions. The scalar intent is therefore a temporal coordination interface rather than a compressed hand pose. Across four simulated tasks, three VLA backbones, and a physical arm--hand platform, InDex preserves the VLA's reaching competence while markedly improving conversion from approach to stable grasp and task completion. Ablations isolate complementary roles: intent aligns the contact transition, whereas diffusion represents the multiple hand trajectories compatible with the same task-space plan. These results identify post-reach arm--hand coordination, rather than object localization alone, as the principal bottleneck in adapting parallel-gripper VLAs to dexterous manipulation.

📄 PDF Abstract BibTeX arXiv:2606.12109

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot GeneralizationSpatial Reasoning

Similar Papers 제목 키워드 기반

Towards Adaptive, Scalable, and Robust Coordination of LLM Agents: A Dynamic Ad-Hoc Networking Perspective

2026-02-08 · Rui Li, Zeyu Zhang, Xiaohe Bo, Quanyu Dai 외 arxiv

Multi-agent architectures built on large language models (LLMs) have demonstrated the potential to realize swarm intelligence through well-crafted collaboration. However, the substantial burden of manual orchestration in…

Parsing Coordination for Spoken Language Understanding

2018-10-26 · Sanchit Agarwal, Rahul Goel, Tagyoung Chung, Abhishek Sethi 외

Typical spoken language understanding systems provide narrow semantic parses using a domain-specific ontology. The parses contain intents and slots that are directly consumed by downstream domain applications. In this wo…

Spoken Language Understanding

Implicit Multiagent Coordination at Unsignalized Intersections via Multimodal Inference Enabled by Topological Braids

2020-04-10 · Christoforos Mavrogiannis, Jonathan A. DeCastro, Siddhartha S. Srinivasa

We focus on navigation among rational, non-communicating agents at unsignalized street intersections. Following collision-free motion under such settings demands nuanced implicit coordination among agents. Often, the str…

Decision MakingTrajectory Prediction

Emergent Coordination through Game-Induced Nonlinear Opinion Dynamics

2023-04-05 · Haimin Hu, Kensuke Nakamura, Kai-Chieh Hsu, Naomi Ehrich Leonard 외

We present a multi-agent decision-making framework for the emergent coordination of autonomous agents whose intents are initially undecided. Dynamic non-cooperative games have been used to encode multi-agent interaction,…

Decision Making

Importance of Intent-Sharing for V2X-based Maneuver Coordination

2026-06-23 · Rafael Molina-Masegosa, Sergei S. Avedisov, Miguel Sepulcre, Javier Gozalvez 외 arxiv

This paper examines the critical role of intent-sharing in enabling effective maneuver coordination for connected and automated vehicles (CAVs). Successful maneuver coordinations require vehicles to accurately know other…

Trajectory Prediction