paper-with-me

홈 › Papers

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation

2026-02-07 · Changhua Xu, En Yu, Junyu Xuan, Jie Lu arxiv

Vision--Language--Action (VLA) models bridge multimodal reasoning with physical control, but adapting them to new tasks with scarce demonstrations remains unreliable. While fine-tuned VLA policies often produce semantically plausible trajectories, failures often arise from unresolved geometric ambiguities, where near-miss actions lead to divergent execution outcomes under limited supervision. We study few-shot VLA adaptation from a \emph{generation--selection} perspective and propose a novel framework \textbf{VGAS} (\textbf{V}alue-\textbf{G}uided \textbf{A}ction-chunk \textbf{S}election). It performs inference-time best-of-$N$ selection to identify action chunks that are both semantically faithful and geometrically precise. Specifically, \textbf{VGAS} employs a finetuned VLA as a high-recall proposal generator and introduces the \textrm{Q-Chunk-Former}, a geometrically grounded Transformer critic to resolve fine-grained geometric ambiguities. In addition, we propose \textit{Explicit Geometric Regularization} (\texttt{EGR}), which shapes a discriminative value landscape to preserve action ranking resolution among near-miss candidates while mitigating value instability under scarce supervision. Experiments and theoretical analysis demonstrate that \textbf{VGAS} consistently improves success rates and robustness under limited demonstrations and distribution shifts. Our code is available at https://github.com/Jyugo-15/VGAS.

📄 PDF Abstract BibTeX arXiv:2602.07399

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Reasoning

Similar Papers 제목 키워드 기반

Chunk-Guided Q-Learning

2026-03-14 · Gwanwoo Song, Kwanyoung Park, Youngwoon Lee arxiv

In offline reinforcement learning (RL), single-step temporal-difference (TD) learning can suffer from bootstrapping error accumulation over long horizons. Action-chunked TD methods mitigate this by backing up over multip…

Reinforcement Learning

Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning

2026-06-01 · Atoosa Chegini, Soheil Feizi arxiv

Selecting the best response from multiple small-model samples using a stronger scorer is a simple inference-time strategy, but fails when the small model has already committed to incorrect reasoning paths. PRM guided sea…

Mathematical Reasoning

Graph-Guided Concept Selection for Efficient Retrieval-Augmented Generation

2025-10-28 · Ziyu Liu, Yijing Liu, Jianfei Yuan, Minzhi Yan 외 arxiv

Graph-based RAG constructs a knowledge graph (KG) from text chunks to enhance retrieval in Large Language Model (LLM)-based question answering. It is especially beneficial in domains such as biomedicine, law, and politic…

Question Answering

Decoupled Q-Chunking

2025-12-11 · Qiyang Li, Seohong Park, Sergey Levine arxiv

Temporal-difference (TD) methods learn state and action values efficiently by bootstrapping from their own future value predictions, but such a self-bootstrapping mechanism is prone to bootstrapping bias, where the error…

Geometry Guided Self-Consistency for Physical AI

2026-05-09 · Yinwei Dai, Zhuofu Chen, Lijie Yang, Ravi Netravali arxiv

State-of-the-art physical AI models generate a chunk of actions per inference through diffusion or flow matching, iteratively refining an initial noise sample into an action trajectory. Because this inference process is …