paper-with-me

홈 › Papers

Sword: Style-Robust World Models as Simulators via Dynamic Latent Bootstrapping for VLA Policy Post-Training

2026-05-08 · Jiaxuan Gao, Yongjian Guo, Zhong Guan, Wen Huang, Wanlun Ma, Xi Xiao, Junwu Xiong, Sheng Wen arxiv

The integration of Vision-Language-Action (VLA) models with World Models has gained increasing attention. One representative approach treats learned World Models as generative simulators, enabling policy optimization entirely within "imagination." However, when deployed as simulators for specific environments such as the LIBERO benchmark, existing World Models often suffer from poor generalization and long-horizon error accumulation. During closed-loop rollouts, these models are highly sensitive to initial-state perturbations; minor changes in color, illumination, and other visual factors can trigger cascading hallucinations, leading to severe blurriness or overexposure. Moreover, long-horizon error accumulation further degrades the quality and fidelity of predicted future states. These issues limit the reliability of World Models as simulators. To mitigate these problems, we propose Sword, a robust World Model framework. Our method introduces Structure-Guided Style Augmentation to disentangle the visual textures of interactive environments from task-relevant dynamics, thereby improving generalization. We further propose Dynamic Latent Bootstrapping, which maintains consistency between training and inference while keeping memory consumption low. Extensive experiments on the LIBERO benchmark show that our method significantly outperforms the baseline WoVR in terms of generalization, generation quality, robustness, fidelity, and the success rate of reinforcement-learning post-training for VLA models.

📄 PDF Abstract BibTeX arXiv:2605.07288

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Dr.Fill: Crosswords and an Implemented Solver for Singly Weighted CSPs

2014-01-18 · Matthew L. Ginsberg

We describe Dr.Fill, a program that solves American-style crossword puzzles. From a technical perspective, Dr.Fill works by converting crosswords to weighted CSPs, and then using a variety of novel techniques to find a s…

The WebCrow French Crossword Solver

2023-11-27 · Giovanni Angelini, Marco Ernandes, Tommaso laquinta, Caroline Stehlé 외

Crossword puzzles are one of the most popular word games, played in different languages all across the world, where riddle style can vary significantly from one country to another. Automated crossword resolution is chall…

Knowledge Graphs

TransWorldNG: Traffic Simulation via Foundation Model

2023-05-25 · Ding Wang, Xuhong Wang, Liang Chen, Shengyue Yao 외

Traffic simulation is a crucial tool for transportation decision-making and policy development. However, achieving realistic simulations in the face of the high dimensionality and heterogeneity of traffic environments is…

Decision MakingManagementmodel

Harnessing LLMs for Educational Content-Driven Italian Crossword Generation

2024-11-25 · Kamyar Zeinalipour, Achille Fusco, Asya Zanollo, Marco Maggini 외

In this work, we unveil a novel tool for generating Italian crossword puzzles from text, utilizing advanced language models such as GPT-4o, Mistral-7B-Instruct-v0.3, and Llama3-8b-Instruct. Crafted specifically for educa…

RiDDLE: Reversible and Diversified De-identification with Latent Encryptor

2023-03-09 · CVPR 2023 1 · Dongze Li, Wei Wang, Kang Zhao, Jing Dong 외

This work presents RiDDLE, short for Reversible and Diversified De-identification with Latent Encryptor, to protect the identity information of people from being misused. Built upon a pre-learned StyleGAN2 generator, RiD…

De-identificationDiversity