paper-with-me

Papers

ECHO: Continuous Hierarchical Memory for Vision-Language-Action Models

2026-05-09 · Yanbin Hu, Jin Cui, Jiayi Lu, Ruixuan Yang, Jun Ye, Boran Zhao, Xingyu Chen, Xuguang Lan, Pengju Ren arxiv

Memory capacity is a critical factor determining the performance of Vision-Language-Action (VLA) models in long-horizon manipulation tasks. Existing memory-augmented architectures primarily rely on linear or flat storage, lacking structural priors for manipulation categories and hierarchical organization. This deficiency hinders efficient experience retrieval and limits generalization to unseen long-horizon task compositions. Inspired by the hierarchical organization of human experience, we propose ECHO (Experience Consolidation and Hierarchical Organization), a novel memory framework operating within a Continuous Hierarchical Space. By employing a hyperbolic autoencoder, ECHO maps VLA hidden states into this space. Leveraging hyperbolic metrics and entailment constraint mechanisms, experience vectors are organized into a semantic memory tree that supports efficient top-down retrieval. In parallel, a background consolidation mechanism continuously refines the memory tree through geometric interpolation and structural splitting, supporting virtual memory synthesis in the continuous space. We integrate ECHO into the $π_0$ foundation model. Evaluations on LIBERO and preliminary real-world experiments demonstrate the effectiveness of our approach, notably achieving a 12.8% absolute improvement in execution success rate over the $π_0$ baseline on LIBERO-Long, while improving compositional generalization on cross-suite unseen long-horizon tasks.

📄 PDF Abstract BibTeX arXiv:2605.10993

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Decoding the Echoes of Vision from fMRI: Memory Disentangling for Past Semantic Information

2024-09-30 · Runze Xia, Congchi Yin, Piji Li

The human visual system is capable of processing continuous streams of visual information, but how the brain encodes and retrieves recent visual memories during continuous visual processing remains unexplored. This study…

Contrastive Learning

EchoVLA: Synergistic Declarative Memory for VLA-Driven Mobile Manipulation

2025-11-22 · Min Lin, Xiwen Liang, Bingqian Lin, Liu Jingzhi 외 arxiv

Recent progress in Vision-Language-Action (VLA) models has enabled embodied agents to interpret multimodal instructions and perform complex tasks. However, existing VLAs are mostly confined to short-horizon, table-top ma…

Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation

2026-05-15 · Mingqiang Wu, Weilun Feng, Zhefeng Zhang, Haotong Qin 외 arxiv

Autoregressive video diffusion models enable open-ended generation through local attention and KV caching. However, existing training-free long-video optimization methods mainly focus on stable extension under a single p…

Video Generation

Echo state networks are universal

2018-06-03 · Lyudmila Grigoryeva, Juan-Pablo Ortega

This paper shows that echo state networks are universal uniform approximants in the context of discrete-time fading memory filters with uniformly bounded inputs defined on negative infinite times. This result guarantees …

valid

The Echo Amplifies the Knowledge: Somatic Marker Analogues in Language Models via Emotion Vector Re-Injection

2026-05-09 · Jared Glover arxiv

Current language model memory systems store what happened but not how it felt. This distinction -- between semantic memory (knowing about a past event) and episodic memory (re-experiencing it) -- was identified by Tulvin…