paper-with-me

홈 › Papers

Learn Where Outcomes Diverge: Efficient VLA RL via Probabilistic Chunk Masking

2026-05-15 · Vaidehi Bagaria, Nikshep Grampurohit, Pulkit Verma arxiv

Reinforcement learning (RL) allows vision-language-action (VLA) policies to generalize beyond their training distribution by optimizing directly for task success, but post-training is computationally expensive. A natural response has been to speed rollout collection through faster simulators and world models. In GRPO-based VLA RL, we find that the dominant cost lies elsewhere: gradient computation accounts for approximately 78% of wall-clock time per step in our runs, while rollout collection accounts for only 21%. Gradient cost dominates because much of this computation is spent on phases that contribute little to learning. GRPO's learning signal is driven by advantage variance: only phases where successful and failed rollouts diverge produce learning signal. However, GRPO assigns the same advantage to every chunk in a rollout. As a result, actor-update compute is spent uniformly across the trajectory, including phases the policy already handles after pre-training and supervised fine-tuning. This paper presents Probabilistic Chunk Masking (PCM), a drop-in modification to GRPO that allocates gradient computation to a small, probabilistically selected subset of chunks per trajectory. PCM scores semantic phases using success-failure action variance, a rollout-derived proxy for per-phase gradient variance, and samples a fixed chunk budget with online-updated phase-level keep probabilities. We formalize per-phase gradient variance as the quantity determines where gradient computation is useful and show that success-failure action variance provides a measurable proxy for it. PCM requires no reward model or learned critic. On three LIBERO benchmarks, PCM matches the final success rate of standard GRPO while achieving 2.38 times wall-clock speedup, 4.8 times faster gradient updates, and 60% lower peak activation memory, while backpropagating through fewer than 20% of trajectory chunks.

📄 PDF Abstract BibTeX arXiv:2605.16154

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation

2026-02-07 · Changhua Xu, En Yu, Junyu Xuan, Jie Lu arxiv

Vision--Language--Action (VLA) models bridge multimodal reasoning with physical control, but adapting them to new tasks with scarce demonstrations remains unreliable. While fine-tuned VLA policies often produce semantica…

Multimodal Reasoning

MEPIC: Memory Efficient Position Independent Caching for LLM Serving

2025-12-18 · Qian Wang, Zahra Yousefijamarani, Morgan Lindsay Heisler, Rongzhi Gu 외 arxiv

Modern LLM applications such as deep-research assistants, coding agents, and Retrieval-Augmented Generation (RAG) systems, repeatedly process long prompt histories containing shared document or code chunks, creating sign…

MoM: Mixtures of Scenario-Aware Document Memories for Retrieval-Augmented Generation Systems

2025-10-16 · Jihao Zhao, Zhiyuan Ji, Simin Niu, Hanyu Wang 외 arxiv

The traditional RAG paradigm, which typically engages in the comprehension of relevant text chunks in response to received queries, inherently restricts both the depth of knowledge internalization and reasoning capabilit…

Soft-NBCE: Entropy-Weighted Chunk Fusion for Long-Context

2026-05-31 · Shihao Ji, Mingyu Li, Zihui Song arxiv

The quadratic complexity of self-attention remains a bottleneck for Large Language Models (LLMs) processing ultra-long contexts. The Naive Bayes Cognitive Engine (NBCE) parallelizes long-context inference by chunking doc…

Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning

2026-06-24 · Jaeyong Ko, Pilsung Kang, Yukyung Lee arxiv

Large language models (LLMs) reach high accuracy in mathematical reasoning, but individual traces on the same problem diverge; some arrive at the correct answer while others fail. Prior work analyzes failure at the step,…

Mathematical Reasoning