Does Semantic Noise Initialization Transfer from Images to Videos? A Paired Diagnostic Study
Semantic noise initialization has been reported to improve robustness and controllability in image diffusion models. Whether these gains transfer to text-to-video (T2V) generation remains unclear, since temporal coupling can introduce extra degrees of freedom and instability. We benchmark semantic noise initialization against standard Gaussian noise using a frozen VideoCrafter-style T2V diffusion backbone and VBench on 100 prompts. Using prompt-level paired tests with bootstrap confidence intervals and a sign-flip permutation test, we observe a small positive trend on temporal-related dimensions; however, the 95 percent confidence interval includes zero (p ~ 0.17) and the overall score remains on par with the baseline. To understand this outcome, we analyze the induced perturbations in noise space and find patterns consistent with weak or unstable signal. We recommend prompt-level paired evaluation and noise-space diagnostics as standard practice when studying initialization schemes for T2V diffusion.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
SELECT: SELEctive Context Transfer for Class-Incremental Semantic Segmentation
Class-Incremental Semantic Segmentation (CISS) is fundamentally challenged by catastrophic forgetting and background shift, where learning new concepts degrades performance on previously seen classes. While existing meth…
Semantic SegmentationParameter-Transferred Wasserstein Generative Adversarial Network (PT-WGAN) for Low-Dose PET Image Denoising
Due to the widespread use of positron emission tomography (PET) in clinical practice, the potential risk of PET-associated radiation dose to patients needs to be minimized. However, with the reduction in the radiation do…
DenoisingDiagnosticGenerative Adversarial NetworkImage Denoising+1Dual-level Semantic Transfer Deep Hashing for Efficient Social Image Retrieval
Social network stores and disseminates a tremendous amount of user shared images. Deep hashing is an efficient indexing technique to support large-scale social image retrieval, due to its deep representation capability, …
Deep HashingImage RetrievalRepresentation LearningRetrievalThe Lottery Ticket Hypothesis in Denoising: Towards Semantic-Driven Initialization
Text-to-image diffusion models allow users control over the content of generated images. Still, text-to-image generation occasionally leads to generation failure requiring users to generate dozens of images under the sam…
DenoisingImage GenerationText to Image GenerationText-to-Image GenerationSpeech-Image Semantic Alignment Does Not Depend on Any Prior Classification Tasks
Semantically-aligned $(speech, image)$ datasets can be used to explore "visually-grounded speech". In a majority of existing investigations, features of an image signal are extracted using neural networks "pre-trained" o…
General ClassificationRetrievalTransfer Learning