paper-with-me

Papers

AnchorEdit: Maintaining Temporal Consistency in Multi-turn Image Editing via Causal Memory

2026-06-10 · Hang Xu, Xiaoxiao Ma, Guohui Zhang, Yu Hu, Siming Fu, Jie Huang, Lin Song, Haoyang Huang, Nan Duan, Feng Zhao arxiv

Multi-turn image editing is essential for iterative design, yet current models often struggle with identity drift and error accumulation over successive steps. While existing research leverages video priors for consistency, their reliance on bidirectional attention is fundamentally misaligned with the causal, sequential nature of interactive editing. In this paper, we propose AnchorEdit, the first autoregressive (AR) diffusion-based framework designed specifically for high-resolution, long-term multi-turn editing. AnchorEdit bridges the gap between video priors and causal inference through a three-stage training curriculum: identity-preserving sing-turn pretraining, causal AR forcing fine-tuning with a novel self-rollout strategy to mitigate exposure bias, and consistency distillation for efficient 4-step generation. During inference, we introduce a memory mechanism to anchor the initial subject identity and ensure stable extrapolation across extended editing trajectories. To evaluate performance, we provide a new high-resolution multi-turn editing benchmark designed to stress-test long-horizon stability. Extensive experiments demonstrate that AnchorEdit achieves state-of-the-art results, maintaining exceptional subject fidelity and instruction following even over 10+ interaction rounds.

📄 PDF Abstract BibTeX arXiv:2606.11751

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction FollowingCausal InferenceImage Editing

Similar Papers 제목 키워드 기반

Session Risk Memory (SRM): Temporal Authorization for Deterministic Pre-Execution Safety Gates

2026-03-22 · Florin Adrian Chitan arxiv

Deterministic pre-execution safety gates evaluate whether individual agent actions are compatible with their assigned roles. While effective at per-action authorization, these systems are structurally blind to distribute…

Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing

2026-01-22 · Dohun Lee, Chun-Hao Paul Huang, Xuelin Chen, Jong Chul Ye 외 arxiv

Video-to-video diffusion models achieve impressive single-turn editing performance, but practical editing workflows are inherently iterative. When edits are applied sequentially, existing models treat each turn independe…

Novel View Synthesis

Counterfactual-Consistency Prompting for Relative Temporal Understanding in Large Language Models

2025-02-17 · Jongho Kim, Seung-won Hwang

Despite the advanced capabilities of large language models (LLMs), their temporal reasoning ability remains underdeveloped. Prior works have highlighted this limitation, particularly in maintaining temporal consistency w…

counterfactual

Depth-Guided Metric-Aware Temporal Consistency for Monocular Video Human Mesh Recovery

2026-02-04 · Jiaxin Cen, Xudong Mao, Guanghui Yue, Wei Zhou 외 arxiv

Monocular video human mesh recovery faces fundamental challenges in maintaining metric consistency and temporal stability due to inherent depth ambiguities and scale uncertainties. While existing methods rely primarily o…

Computational EfficiencyHuman Mesh Recovery

NLUT: Neural-based 3D Lookup Tables for Video Photorealistic Style Transfer

2023-03-16 · Yaosen Chen, Han Yang, Yuexin Yang, Yuegen Liu 외

Video photorealistic style transfer is desired to generate videos with a similar photorealistic style to the style image while maintaining temporal consistency. However, existing methods obtain stylized video sequences b…

8kStyle Transfer