paper-with-me

홈 › Papers

Frame by Familiar Frame: Understanding Replication in Video Diffusion Models

2024-03-28 · Aimon Rahman, Malsha V. Perera, Vishal M. Patel

Building on the momentum of image generation diffusion models, there is an increasing interest in video-based diffusion models. However, video generation poses greater challenges due to its higher-dimensional nature, the scarcity of training data, and the complex spatiotemporal relationships involved. Image generation models, due to their extensive data requirements, have already strained computational resources to their limits. There have been instances of these models reproducing elements from the training samples, leading to concerns and even legal disputes over sample replication. Video diffusion models, which operate with even more constrained datasets and are tasked with generating both spatial and temporal content, may be more prone to replicating samples from their training sets. Compounding the issue, these models are often evaluated using metrics that inadvertently reward replication. In our paper, we present a systematic investigation into the phenomenon of sample replication in video diffusion models. We scrutinize various recent diffusion models for video synthesis, assessing their tendency to replicate spatial and temporal content in both unconditional and conditional generation scenarios. Our study identifies strategies that are less likely to lead to replication. Furthermore, we propose new evaluation strategies that take replication into account, offering a more accurate measure of a model's ability to generate the original content.

📄 PDF Abstract BibTeX arXiv:2403.19593

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Replication in Visual Diffusion Models: A Survey and Outlook

2024-07-07 · Wenhao Wang, Yifan Sun, Zongxin Yang, Zhengdong Hu 외

Visual diffusion models have revolutionized the field of creative AI, producing high-quality and diverse content. However, they inevitably memorize training images or videos, subsequently replicating their concepts, cont…

BenchmarkingSurvey

The Persistence of Cultural Memory: Investigating Multimodal Iconicity in Diffusion Models

2025-11-14 · Maria-Teresa De Rosa Palmini, Eva Cetinic arxiv

The ambiguity between generalization and memorization in TTI diffusion models becomes pronounced when prompts invoke culturally shared visual references, a phenomenon we term multimodal iconicity. These are instances in …

Image Matching

TVBench: Redesigning Video-Language Evaluation

2024-10-10 · Daniel Cores, Michael Dorkenwald, Manuel Mucientes, Cees G. M. Snoek 외

Large language models have demonstrated impressive performance when integrated with vision models even enabling video understanding. However, evaluating these video models presents its own unique challenges, for which se…

Multiple-choiceOpen-Ended Question AnsweringQuestion AnsweringVideo Understanding+2

Towards Open-Vocabulary Video Semantic Segmentation

2024-12-12 · Xinhao Li, Yun Liu, Guolei Sun, Min Wu 외

Semantic segmentation in videos has been a focal point of recent research. However, existing models encounter challenges when faced with unfamiliar categories. To address this, we introduce the Open Vocabulary Video Sema…

SegmentationSemantic SegmentationVideo Semantic SegmentationZero-shot Generalization

Learning to Alleviate Familiarity Bias in Video Recommendation

2026-02-08 · Zheng Ren, Yi Wu, Jianan Lu, Acar Ary 외 arxiv

Modern video recommendation systems aim to optimize user engagement and platform objectives, yet often face structural exposure imbalances caused by behavioral biases. In this work, we focus on the post-ranking stage and…

Recommendation Systems