paper-with-me

홈 › Papers

Tuning-free Visual Effect Transfer across Videos

2026-01-12 · Maxwell Jones, Rameen Abdal, Or Patashnik, Ruslan Salakhutdinov, Sergey Tulyakov, Jun-Yan Zhu, Kuan-Chieh Jackson Wang arxiv

We present RefVFX, a new framework that transfers complex temporal effects from a reference video onto a target video or image in a feed-forward manner. While existing methods excel at prompt-based or keyframe-conditioned editing, they struggle with dynamic temporal effects such as dynamic lighting changes or character transformations, which are difficult to describe via text or static conditions. Transferring a video effect is challenging, as the model must integrate the new temporal dynamics with the input video's existing motion and appearance. % To address this, we introduce a large-scale dataset of triplets, where each triplet consists of a reference effect video, an input image or video, and a corresponding output video depicting the transferred effect. Creating this data is non-trivial, especially the video-to-video effect triplets, which do not exist naturally. To generate these, we propose a scalable automated pipeline that creates high-quality paired videos designed to preserve the input's motion and structure while transforming it based on some fixed, repeatable effect. We then augment this data with image-to-video effects derived from LoRA adapters and code-based temporal effects generated through programmatic composition. Building on our new dataset, we train our reference-conditioned model using recent text-to-video backbones. Experimental results demonstrate that RefVFX produces visually consistent and temporally coherent edits, generalizes across unseen effect categories, and outperforms prompt-only baselines in both quantitative metrics and human preference. See our website at https://snap-research.github.io/RefVFX/

📄 PDF Abstract BibTeX arXiv:2601.07833

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An Analysis of Layer-Freezing Strategies for Enhanced Transfer Learning in YOLO Architectures

2025-09-05 · Andrzej D. Dobrzycki, Ana M. Bernardos, José R. Casar arxiv

The You Only Look Once (YOLO) architecture is crucial for real-time object detection. However, deploying it in resource-constrained environments such as unmanned aerial vehicles (UAVs) requires efficient transfer learnin…

Real-Time Object DetectionTransfer Learning

When Fine-Tuning Changes the Evidence: Architecture-Dependent Semantic Drift in Chest X-Ray Explanations

2026-04-09 · Kabilan Elangovan, Daniel Ting arxiv

Transfer learning followed by fine-tuning is widely adopted in medical image classification due to consistent gains in diagnostic performance. However, in multi-class settings with overlapping visual features, improvemen…

Medical Image ClassificationTransfer LearningVisual Reasoning

NVSMask3D: Hard Visual Prompting with Camera Pose Interpolation for 3D Open Vocabulary Instance Segmentation

2025-04-20 · Junyuan Fang, Zihan Wang, Yejun Zhang, Shuzhe Wang 외

Vision-language models (VLMs) have demonstrated impressive zero-shot transfer capabilities in image-level visual perception tasks. However, they fall short in 3D instance-level segmentation tasks that require accurate lo…

3D Instance Segmentation3D Open-Vocabulary Instance SegmentationDescriptiveInstance Segmentation+2

Decouple before Align: Visual Disentanglement Enhances Prompt Tuning

2025-08-01 · Fei Zhang, Tianfei Zhou, Jiangchao Yao, Ya Zhang 외 arxiv

Prompt tuning (PT), as an emerging resource-efficient fine-tuning paradigm, has showcased remarkable effectiveness in improving the task-specific transferability of vision-language models. This paper delves into a previo…

Few-Shot Learning

Consistent Story Generation with Asymmetry Zigzag Sampling

2025-06-11 · Mingxiao Li, Mang Ning, Marie-Francine Moens

Text-to-image generation models have made significant progress in producing high-quality images from textual descriptions, yet they continue to struggle with maintaining subject consistency across multiple images, a fund…

Image GenerationStory GenerationStory VisualizationText to Image Generation+2