paper-with-me

홈 › Papers

VidCLearn: A Continual Learning Approach for Text-to-Video Generation

2025-09-21 · Luca Zanchetta, Lorenzo Papa, Luca Maiano, Irene Amerini arxiv

Text-to-video generation is an emerging field in generative AI, enabling the creation of realistic, semantically accurate videos from text prompts. While current models achieve impressive visual quality and alignment with input text, they typically rely on static knowledge, making it difficult to incorporate new data without retraining from scratch. To address this limitation, we propose VidCLearn, a continual learning framework for diffusion-based text-to-video generation. VidCLearn features a student-teacher architecture where the student model is incrementally updated with new text-video pairs, and the teacher model helps preserve previously learned knowledge through generative replay. Additionally, we introduce a novel temporal consistency loss to enhance motion smoothness and a video retrieval module to provide structural guidance at inference. Our architecture is also designed to be more computationally efficient than existing models while retaining satisfactory generation performance. Experimental results show VidCLearn's superiority over baseline methods in terms of visual quality, semantic alignment, and temporal coherence.

📄 PDF Abstract BibTeX arXiv:2509.16956

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video GenerationContinual LearningVideo Retrieval

Similar Papers 제목 키워드 기반

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement

2024-12-25 · Zhefan Rao, Liya Ji, Yazhou Xing, Runtao Liu 외

Text-to-video (T2V) generation has gained significant attention recently. However, the costs of training a T2V model from scratch remain persistently high, and there is considerable room for improving the generation perf…

Bring Your Dreams to Life: Continual Text-to-Video Customization

2025-12-05 · Jiahua Dong, Xudong Wang, Wenqi Liang, Zongyan Han 외 arxiv

Customized text-to-video generation (CTVG) has recently witnessed great progress in generating tailored videos from user-specific text. However, most CTVG methods assume that personalized concepts remain static and do no…

Text-to-Video GenerationNoise Estimation

StableFusion: Continual Video Retrieval via Frame Adaptation

2025-03-13 · Zecheng Zhao, Zhi Chen, Zi Huang, Shazia Sadiq 외

Text-to-Video Retrieval (TVR) aims to match videos with corresponding textual queries, yet the continual influx of new video content poses a significant challenge for maintaining system performance over time. In this wor…

Continual LearningMixture-of-ExpertsRetrievalText to Video Retrieval+1

ViLCo-Bench: VIdeo Language COntinual learning Benchmark

2024-06-19 · Tianqi Tang, Shohreh Deldari, Hao Xue, Celso de Melo 외

Video language continual learning involves continuously adapting to information from video and text inputs, enhancing a model's ability to handle new tasks while retaining prior knowledge. This field is a relatively unde…

Continual LearningSelf-Supervised Learning

HyperTokens: Controlling Token Dynamics for Continual Video-Language Understanding

2026-03-02 · Toan Nguyen, Yang Liu, Celso De Melo, Flora D. Salim arxiv

Continual VideoQA with multimodal LLMs is hindered by interference between tasks and the prohibitive cost of storing task-specific prompts. We introduce HyperTokens, a transformer-based token generator that produces fine…