paper-with-me

Papers

FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment

2026-02-19 · Han Zhao, Jingbo Wang, Wenxuan Song, Shuai Chen, Yang Liu, Yan Wang, Haoang Li, Donglin Wang arxiv

Enabling VLA models to predict environmental dynamics, known as world modeling, has been recognized as essential for improving robotic reasoning and generalization. However, current approaches face two main issues: 1. The training objective forces models to over-emphasize pixel-level reconstruction, which constrains semantic learning and generalization 2. Reliance on predicted future observations during inference often leads to error accumulation. To address these challenges, we introduce Future Representation Alignment via Parallel Progressive Expansion (FRAPPE). Our method adopts a two-stage fine-tuning strategy: In the mid-training phase, the model learns to predict the latent representations of future observations; In the post-training phase, we expand the computational workload in parallel and align the representation simultaneously with multiple different visual foundation models. By significantly improving fine-tuning efficiency and reducing dependence on action-annotated data, FRAPPE provides a scalable and data-efficient pathway to enhance world-awareness in generalist robotic policies. Experiments on the RoboTwin benchmark and real-world tasks demonstrate that FRAPPE outperforms state-of-the-art approaches and shows strong generalization in long-horizon and unseen scenarios.

📄 PDF Abstract BibTeX arXiv:2602.17259

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Frappe: Understanding the Usage and Perception of Mobile App Recommendations In-The-Wild

2015-05-12 · Baltrunas Linas, Church Karen, Karatzoglou Alexandros, Oliver Nuria

This paper describes a real world deployment of a context-aware mobile app recommender system (RS) called Frappe. Utilizing a hybrid-approach, we conducted a large-scale app market deployment with 1000 Android users comb…

Recommendation Systems

FRAPPE: $\underline{\text{F}}$ast $\underline{\text{Ra}}$nk $\underline{\text{App}}$roximation with $\underline{\text{E}}$xplainable Features for Tensors

2022-06-19 · William Shiao, Evangelos E. Papalexakis

Tensor decompositions have proven to be effective in analyzing the structure of multidimensional data. However, most of these methods require a key parameter: the number of desired components. In the case of the CANDECOM…

X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model

2025-10-11 · Jinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu 외 arxiv

Successful generalist Vision-Language-Action (VLA) models rely on effective training across diverse robotic platforms with large-scale, cross-embodiment, heterogeneous datasets. To facilitate and leverage the heterogenei…

FraPPE: Fast and Efficient Preference-based Pure Exploration

2025-08-22 · Udvas Das, Apurv Shukla, Debabrota Basu arxiv

Preference-based Pure Exploration (PrePEx) aims to identify with a given confidence level the set of Pareto optimal arms in a vector-valued (aka multi-objective) bandit, where the reward vectors are ordered via a (given)…

Hydra-0: Action Flow for Generalist World Modeling and Control

2026-08-18 · Hongyu Li, Bowen Wen, Xinghao Zhu, Yixuan Wang 외 arxiv

We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action con…