paper-with-me

홈 › Papers

Learning Long-Context Diffusion Policies via Past-Token Prediction

2025-05-14 · Marcel Torne, Andy Tang, Yuejiang Liu, Chelsea Finn

Reasoning over long sequences of observations and actions is essential for many robotic tasks. Yet, learning effective long-context policies from demonstrations remains challenging. As context length increases, training becomes increasingly expensive due to rising memory demands, and policy performance often degrades as a result of spurious correlations. Recent methods typically sidestep these issues by truncating context length, discarding historical information that may be critical for subsequent decisions. In this paper, we propose an alternative approach that explicitly regularizes the retention of past information. We first revisit the copycat problem in imitation learning and identify an opposite challenge in recent diffusion policies: rather than over-relying on prior actions, they often fail to capture essential dependencies between past and future actions. To address this, we introduce Past-Token Prediction (PTP), an auxiliary task in which the policy learns to predict past action tokens alongside future ones. This regularization significantly improves temporal modeling in the policy head, with minimal reliance on visual representations. Building on this observation, we further introduce a multistage training strategy: pre-train the visual encoder with short contexts, and fine-tune the policy head using cached long-context embeddings. This strategy preserves the benefits of PTP while greatly reducing memory and computational overhead. Finally, we extend PTP into a self-verification mechanism at test time, enabling the policy to score and select candidates consistent with past actions during inference. Experiments across four real-world and six simulated tasks demonstrate that our proposed method improves the performance of long-context diffusion policies by 3x and accelerates policy training by more than 10x.

📄 PDF Abstract BibTeX arXiv:2505.09561

Code (0)

등록된 구현이 없습니다.

Tasks

Imitation Learning

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing

2026-02-02 · Lingkun Long, Yushi Huang, Shihao Bai, Ruihao Gong 외 arxiv

Diffusion Large Language Models (dLLMs) deliver strong long-context processing capability in a non-autoregressive decoding paradigm. However, the considerable computational cost of bidirectional full attention limits the…

ChronoFlow-Policy: Unifying Past-Current-Future Interaction Flow in Visuomotor Policy Learning

2026-06-30 · Bokai Lin, Yifu Xu, Xinyu Zhan, Hongjie Fang 외 arxiv

Visual signals play a crucial role in policy learning by enabling models to capture object motion and interaction dynamics. Just as humans reason about actions using both past experience and anticipated outcomes, effecti…

VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory

2026-03-05 · Yuheng Lei, Zhixuan Liang, Hongyuan Zhang, Ping Luo arxiv

Imitation learning from human demonstrations has achieved significant success in robotic control, yet most visuomotor policies still condition on single-step observations or short-context histories, making them struggle …

Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion

2026-01-29 · Hanmo Chen, Chenghao Xu, Xu Yang, Xuan Chen 외 arxiv

Video generation is pivotal to digital media creation, and recent advances in autoregressive video generation have markedly enhanced the efficiency of real-time video synthesis. However, existing approaches generally rel…

Video Generation

MemRoPE: Training-Free Infinite Video Generation via Evolving Memory Tokens

2026-03-12 · Youngrae Kim, Qixin Hu, C. -C. Jay Kuo, Peter A. Beerel arxiv

Autoregressive diffusion enables real-time frame streaming, yet existing sliding-window caches discard past context, causing fidelity degradation, identity drift, and motion stagnation over long horizons. Current approac…

Video Generation