paper-with-me

홈 › Papers

VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers

2026-05-17 · Yiren Song, Wangzi Yao, Haofan Wang, Mike Zheng Shou arxiv

Video style transfer aims to render videos in a target artistic style while preserving content, structure, and motion. While image stylization has advanced rapidly, video stylization remains challenging due to temporal inconsistency. Most existing methods stylize frames or keyframes and enforce consistency via heuristic temporal propagation, which is brittle under occlusions, disocclusions, and long-term motion, leading to drift and flickering artifacts. We argue that a fundamental bottleneck lies in the lack of large-scale triplet data and a principled training paradigm that jointly models and disentangles style, content, and motion.To address this, we introduce VISTA-1000, a synthetic dataset with 1,000 styles and motion-aligned triplets of style reference, clean video, and stylized video, and propose a diffusion-transformer-based in-context video style transfer framework with a lightweight style adapter for robust style extraction. Extensive experiments demonstrate SOTA performance in style fidelity, temporal consistency, and content preservation.

📄 PDF Abstract BibTeX arXiv:2605.17312

Code (0)

등록된 구현이 없습니다.

Tasks

Style Transfer

Similar Papers 제목 키워드 기반

TeleStyle: Content-Preserving Style Transfer in Images and Videos

2026-01-28 · Shiwen Zhang, Xiaoyan Yang, Bojia Zi, Haibin Huang 외 arxiv

Content-preserving style transfer, generating stylized outputs based on content and style references, remains a significant challenge for Diffusion Transformers (DiTs) due to the inherent entanglement of content and styl…

Continual LearningStyle Transfer

Towards In-Context Tone Style Transfer with A Large-Scale Triplet Dataset

2026-04-17 · Yuhai Deng, Huimin She, Wei Shen, Meng Li 외 arxiv

Tone style transfer for photo retouching aims to adapt the stylistic tone of the reference image to a given content image. However, the lack of high-quality large-scale triplet datasets with stylized ground truth forces …

Photo RetouchingStyle Transfer

OmniStyle: Filtering High Quality Style Transfer Data at Scale

2025-05-20 · CVPR 2025 1 · Ye Wang, Ruiqi Liu, Jiang Lin, Fei Liu 외

In this paper, we introduce OmniStyle-1M, a large-scale paired style transfer dataset comprising over one million content-style-stylized image triplets across 1,000 diverse style categories, each enhanced with textual de…

Style Transfer

Image Style Transfer and Content-Style Disentanglement

2021-11-25 · Sailun Xu, Jiazhi Zhang, Jiamei Liu

We propose a way of learning disentangled content-style representation of image, allowing us to extrapolate images to any style as well as interpolate between any pair of styles. By augmenting data set in a supervised se…

DisentanglementStyle TransferTriplet

Pureformer-VC: Non-parallel One-Shot Voice Conversion with Pure Transformer Blocks and Triplet Discriminative Training

2024-09-03 · Wenhan Yao, Zedong Xing, Xiarun Chen, Jia Liu 외

One-shot voice conversion(VC) aims to change the timbre of any source speech to match that of the target speaker with only one speech sample. Existing style transfer-based VC methods relied on speech representation disen…

DecoderDisentanglementStyle TransferTriplet+1