paper-with-me

홈 › Papers

VALA: Learning Latent Anchors for Training-Free and Temporally Consistent

2025-10-27 · Zhangkai Wu, Xuhui Fan, Zhongyuan Xie, Kaize Shi, Longbing Cao arxiv

Recent advances in training-free video editing have enabled lightweight and precise cross-frame generation by leveraging pre-trained text-to-image diffusion models. However, existing methods often rely on heuristic frame selection to maintain temporal consistency during DDIM inversion, which introduces manual bias and reduces the scalability of end-to-end inference. In this paper, we propose~\textbf{VALA} (\textbf{V}ariational \textbf{A}lignment for \textbf{L}atent \textbf{A}nchors), a variational alignment module that adaptively selects key frames and compresses their latent features into semantic anchors for consistent video editing. To learn meaningful assignments, VALA propose a variational framework with a contrastive learning objective. Therefore, it can transform cross-frame latent representations into compressed latent anchors that preserve both content and temporal coherence. Our method can be fully integrated into training-free text-to-image based video editing models. Extensive experiments on real-world video editing benchmarks show that VALA achieves state-of-the-art performance in inversion fidelity, editing quality, and temporal consistency, while offering improved efficiency over prior methods.

📄 PDF Abstract BibTeX arXiv:2510.22970

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive Learning

Similar Papers 제목 키워드 기반

PVALane: Prior-Guided 3D Lane Detection with View-Agnostic Feature Alignment

2024-03-24 · Proceedings of the AAAI Conference on Artificial Intelligence 2024 3 · Zewen Zheng, Xuemin Zhang, Yongqiang Mou, Xiang Gao 외

Monocular 3D lane detection is essential for a reliable autonomous driving system and has recently been rapidly developing. Existing popular methods mainly employ a predefined 3D anchor for lane detection based on front-…

3D Lane DetectionAutonomous DrivingLane Detection

Neural criticality from effective latent variables

2023-01-02 · Mia C. Morrell, Ilya Nemenman, Audrey J. Sederberg

Observations of power laws in neural activity data have raised the intriguing notion that brains may operate in a critical state. One example of this critical state is "avalanche criticality," which has been observed in …

EEGElectroencephalogram (EEG)

FlowAnchor: Stabilizing the Editing Signal for Inversion-Free Video Editing

2026-04-24 · Ze Chen, Lan Chen, Yuanhang Li, Qi Mao arxiv

We propose FlowAnchor, a training-free framework for stable and efficient inversion-free, flow-based video editing. Inversion-free editing methods have recently shown impressive efficiency and structure preservation in i…

OnlineSplatter: Pose-Free Online 3D Reconstruction for Free-Moving Objects

2025-10-23 · Mark He Huang, Lin Geng Foo, Christian Theobalt, Ying Sun 외 arxiv

Free-moving object reconstruction from monocular video remains challenging, particularly without reliable pose or depth cues and under arbitrary object motion. We introduce OnlineSplatter, a novel online feed-forward fra…

3D Reconstruction

Improving Joint Audio-Video Generation with Cross-Modal Context Learning

2026-03-19 · Bingqi Ma, Linlong Lang, Ming Zhang, Dailan He 외 arxiv

The dual-stream transformer architecture-based joint audio-video generation method has become the dominant paradigm in current research. By incorporating pre-trained video diffusion models and audio diffusion models, alo…

Video Generation