paper-with-me

홈 › Papers

ChordVideo: One-Step, Training-Free, Temporally Consistent Video Editing via Low-Energy Transport

2026-08-01 · Zhiqiang Lao arxiv

One-step text-to-image models enable training-free, inversion-free editing with only 1--2 network function evaluations (NFE), while ChordEdit stabilizes such edits through low-energy smoothing along sampling time. Applied independently to video frames, however, it produces temporal flicker and edit-strength drift. We introduce \textbf{ChordVideo}, which extends the same low-energy principle to video time through shared noise, motion-aligned causal aggregation of per-frame Chord fields, and an optional temporally smoothed proximal correction. We derive a warping-error bound that separates motion bias from stochastic flicker and predicts diminishing returns with larger temporal windows. On TGVE/DAVIS with two one-step backbones, ChordVideo reduces warping error by \textbf{78\%} and flicker by \textbf{49\%}, improves CLIP frame consistency by \textbf{9--10 points}, and increases background PSNR by about \textbf{1.5,dB}, while retaining \textbf{2 NFE/frame}. Compared with seven multi-step editors, it achieves competitive temporal consistency and source preservation using \textbf{10--60$\times$ fewer model steps per clip

📄 PDF Abstract BibTeX arXiv:2608.00769

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models

2026-04-16 · Yanda Li, Yuhan Liu, Zirui Song, Yunchao Wei 외 arxiv

Large audio-language models (LALMs) generalize across speech, sound, and music, but unified decoders can exhibit a \emph{temporal smoothing bias}: transient acoustic cues may be underutilized in favor of temporally smoot…

4-Doodle: Text to 3D Sketches that Move!

2025-10-29 · Hao Chen, Jiaqi Wang, Yonggang Qi, Ke Li 외 arxiv

We present a novel task: text-to-3D sketch animation, which aims to bring freeform sketches to life in dynamic 3D space. Unlike prior works focused on photorealistic content generation, we target sparse, stylized, and vi…

Point CloudsText to 3D

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation

2026-04-29 · Mingji Ge, Qirui Chen, Zeqian Li, Weidi Xie arxiv

Long-term video understanding requires interpreting complex temporal events and reasoning over procedural activities. While instructional video corpora, like HowTo100M, offer rich resources for model training, they prese…

Zero-shot GeneralizationDense Video CaptioningCross-Modal Retrieval

Dynamic Scene Novel View Synthesis via Deferred Spatio-temporal Consistency

2021-09-02 · Beatrix-Emőke Fülöp-Balogh, Eleanor Tursman, James Tompkin, Julie Digne 외

Structure from motion (SfM) enables us to reconstruct a scene via casual capture from cameras at different viewpoints, and novel view synthesis (NVS) allows us to render a captured scene from a new viewpoint. Both are ha…

Novel View Synthesis

Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss

2025-01-13 · Xinyu Zhang, Zicheng Duan, Dong Gong, Lingqiao Liu

In this paper, we address the challenge of generating temporally consistent videos with motion guidance. While many existing methods depend on additional control modules or inference-time fine-tuning, recent studies sugg…

Feature CorrelationVideo Generation