paper-with-me

홈 › Papers

TIDE: Task-Isolated Diffusion for Unified Video Editing and Generation

2026-06-06 · Qi Liu, Gang Yue, Mingyu Yin, Lisai Zhang, Yidi Wu, Yaole Wang, Yaohui Wang, Chang Yao, Jingyuan Chen, Lin Ma arxiv

Recent advances in Diffusion Transformers have driven rapid progress in video generation and editing, yet these capabilities are still handled by separate, task-specific models. Building a unified framework that supports diverse video tasks remains an open challenge: existing unified attempts either require dedicated auxiliary encoders or lack explicit mechanisms to distinguish heterogeneous conditioning tokens, struggling when the number and type of visual conditions vary across tasks. We propose TIDE, a unified framework that integrates instruction-based editing, reference-guided editing, and multi-reference generation. At its core, we introduce per-token task embeddings that assign each input token a task-specific identifier, enabling the model to explicitly disambiguate target, source, and reference tokens. To simultaneously capture high-level semantic understanding and fine-grained structural fidelity, we design a dual-path conditioning scheme that couples a vision-language model with a VAE latent path for complementary signals. We further devise a multi-task progressive training strategy that incrementally introduces tasks of increasing complexity, effectively harmonizing diverse objectives and enabling smooth generalization across heterogeneous task distributions. Extensive experiments on multiple video editing and generation benchmarks demonstrate that TIDE achieves state-of-the-art performance across all evaluated tasks. Our project page is available at https://LittleWork123.github.io/tide.

📄 PDF Abstract BibTeX arXiv:2606.08260

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Trends for isolated amino acids and dipeptides: Conformation, divalent ion binding, and remarkable similarity of binding to calcium and lead

2016-09-14

We derive structural and binding energy trends for twenty amino acids, their dipeptides, and their interactions with the divalent cations Ca$^{2+}$, Ba$^{2+}$, Sr$^{2+}$, Cd$^{2+}$, Pb$^{2+}$, and Hg$^{2+}$. The underlyi…

VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning

2025-10-09 · Minghong Cai, Qiulin Wang, Zongli Ye, Wenze Liu 외 arxiv

Existing controllable video generation methods are typically designed for rigid, task-specific settings, such as first-frame image-to-video, inpainting, or interpolation, treating spatio-temporal control as a set of isol…

Video Generation

UniMoMo: Unified Generative Modeling of 3D Molecules for De Novo Binder Design

2025-03-25 · Xiangzhe Kong, Zishen Zhang, Ziting Zhang, Rui Jiao 외

The design of target-specific molecules such as small molecules, peptides, and antibodies is vital for biological research and drug discovery. Existing generative methods are restricted to single-domain molecules, failin…

Drug DiscoveryLatent Diffusion Model for 3D

Peptide2Mol: A Diffusion Model for Generating Small Molecules as Peptide Mimics for Targeted Protein Binding

2025-11-07 · Xinheng He, Yijia Zhang, Haowei Lin, Xingang Peng 외 arxiv

Structure-based drug design has seen significant advancements with the integration of artificial intelligence (AI), particularly in the generation of hit and lead compounds. However, most AI-driven approaches neglect the…

Graph Neural Network

PepDoRA: A Unified Peptide Language Model via Weight-Decomposed Low-Rank Adaptation

2024-10-28 · Leyao Wang, Rishab Pulugurta, Pranay Vure, Yinuo Zhang 외

Peptide therapeutics, including macrocycles, peptide inhibitors, and bioactive linear peptides, play a crucial role in therapeutic development due to their unique physicochemical properties. However, predicting these pro…

Activity PredictionContrastive LearningLanguage ModelingLanguage Modelling+1