paper-with-me

홈 › Papers

SWEET: Sparse World Modeling with Image Editing for Embodied Task Execution

2026-05-19 · Yiren Song, Yihan Wang, Xiyao Deng, Zhuoran Yan, Mike Zheng Shou arxiv

Visual prediction has emerged as a promising paradigm for embodied control, where future observations are generated and then translated into actions. However, dense video generation is computationally expensive and often unnecessary for many manipulation tasks, whose progress can be summarized by a small number of task-relevant visual states. In this work, we study whether image editing models can serve as sparse visual world models for robot manipulation by predicting task-level future states without dense video rollout. We first conduct a controlled comparison between the video generation model Wan2.2 and the image editing model FLUX-Kontext under the same robotic data setting, and find that image editing produces more reliable task-level keyframes with better visual fidelity and substantially lower inference cost. Motivated by this observation, we propose SWEET, a one-shot sparse visual planning framework that progressively generates a sequence of task-relevant manipulation keyframes through successive image editing, conditioned on language instructions and optional arrow-based spatial guidance. A goal-conditioned diffusion action predictor then converts adjacent imagined keyframes into executable action chunks. To reduce the mismatch between real and edited visual subgoals, we further introduce a mixed-training strategy with filtered edited targets. Experiments on DROID and RoboMimic show that SWEET improves keyframe prediction across seen and unseen scenes and enables a full pipeline from sequential keyframe planning to executable robot actions, suggesting that image editing is a promising and underexplored direction for embodied visual prediction.

📄 PDF Abstract BibTeX arXiv:2605.19319

Code (0)

등록된 구현이 없습니다.

Tasks

Robot ManipulationVideo GenerationImage Editing

Similar Papers 제목 키워드 기반

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

2026-06-17 · Yuyang Zhang, Wenyao Zhang, Zekun Qi, He Zhang 외 arxiv

World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control. However, video-based WAMs face three coupled limitations: dense multi-frame future tokens make inference cos…

Video PredictionVideo GenerationImage Editing

Expediting Contrastive Language-Image Pretraining via Self-distilled Encoders

2023-12-19 · Bumsoo Kim, Jinhyung Kim, Yeonsik Jo, Seung Hwan Kim

Recent advances in vision language pretraining (VLP) have been largely attributed to the large-scale data collected from the web. However, uncurated dataset contains weakly correlated image-text pairs, causing data ineff…

Knowledge Distillation

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models

2025-12-16 · Shufan Li, Jiuxiang Gu, Kangning Liu, Zhe Lin 외 arxiv

Masked Discrete Diffusion Models (MDMs) have achieved strong performance across a wide range of multimodal tasks, including image understanding, generation, and editing. However, their inference speed remains suboptimal …

Text-to-Image GenerationMathematical ReasoningImage Editing

AttnRouter: Per-Category Attention Routing for Training-Free Image Editing on MMDiT

2026-05-02 · Guandong Li, Mengxia Ye arxiv

We study training-free image editing on Qwen-Image-Edit-2511, a 60-block multi-modal diffusion transformer (MMDiT) that concatenates noise and source-image tokens within a single attention stream. We make three contribut…

Image Editing

Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling

2026-05-13 · Xuehai Bai, Yang Shi, Yi-Fan Zhang, Xuanyu Zhu 외 arxiv

Recent image editing models have achieved remarkable progress in instruction following, multimodal understanding, and complex visual editing. However, existing benchmarks often fail to faithfully reflect human judgment, …

Instruction FollowingVisual ReasoningImage Editing