paper-with-me

홈 › Papers

Training-Free Semantic Video Composition via Pre-trained Diffusion Model

2024-01-17 · Jiaqi Guo, Sitong Su, Junchen Zhu, Lianli Gao, Jingkuan Song

The video composition task aims to integrate specified foregrounds and backgrounds from different videos into a harmonious composite. Current approaches, predominantly trained on videos with adjusted foreground color and lighting, struggle to address deep semantic disparities beyond superficial adjustments, such as domain gaps. Therefore, we propose a training-free pipeline employing a pre-trained diffusion model imbued with semantic prior knowledge, which can process composite videos with broader semantic disparities. Specifically, we process the video frames in a cascading manner and handle each frame in two processes with the diffusion model. In the inversion process, we propose Balanced Partial Inversion to obtain generation initial points that balance reversibility and modifiability. Then, in the generation process, we further propose Inter-Frame Augmented attention to augment foreground continuity across frames. Experimental results reveal that our pipeline successfully ensures the visual harmony and inter-frame coherence of the outputs, demonstrating efficacy in managing broader semantic disparities.

📄 PDF Abstract BibTeX arXiv:2401.09195

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Training-Free Semantic Correction for Autoregressive Visual Models

2026-06-21 · Junhao Chen, Chanyu Zhu, Zheqi Lv, Keting Yin 외 arxiv

Autoregressive visual models (AVMs) based on next-scale prediction have emerged as a prominent paradigm for image and video synthesis. However, decomposing the generation process into discrete scales with varying granula…

VideoGEM: Training-free Action Grounding in Videos

2025-03-26 · CVPR 2025 1 · Felix Vogel, Walid Bousselham, Anna Kukleva, Nina Shvetsova 외

Vision-language foundation models have shown impressive capabilities across various zero-shot tasks, including training-free localization and grounding, primarily focusing on localizing objects in images. However, levera…

Video Grounding

NEGATE: Constrained Semantic Guidance for Linguistic Negation in Text-to-Video Diffusion

2026-03-06 · Taewon Kang, Ming C. Lin arxiv

Negation is a fundamental linguistic operator, yet it remains inadequately modeled in diffusion-based generative systems. In this work, we present a formal treatment of linguistic negation in diffusion-based generative m…

Image Generation

MagicComp: Training-free Dual-Phase Refinement for Compositional Video Generation

2025-03-18 · Hongyu Zhang, Yufan Deng, Shenghai Yuan, Peng Jin 외

Text-to-video (T2V) generation has made significant strides with diffusion models. However, existing methods still struggle with accurately binding attributes, determining spatial relationships, and capturing complex act…

DenoisingVideo Generation

VidSeg: Training-free Video Semantic Segmentation based on Diffusion Models

2025-01-01 · CVPR 2025 1 · Qian Wang, Abdelrahman Eldesokey, Mohit Mendiratta, Fangneng Zhan 외

We introduce the first training-free approach for Video Semantic Segmentation (VSS) based on pre-trained diffusion models. A growing research direction attempts to employ diffusion models to perform downstream vision…

SegmentationSemantic SegmentationVideo Semantic Segmentation