paper-with-me

홈 › Papers

VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

2024-08-05 · Zhiyu Tan, Xiaomeng Yang, Luozheng Qin, Hao Li

The quality of video-text pairs fundamentally determines the upper bound of text-to-video models. Currently, the datasets used for training these models suffer from significant shortcomings, including low temporal consistency, poor-quality captions, substandard video quality, and imbalanced data distribution. The prevailing video curation process, which depends on image models for tagging and manual rule-based curation, leads to a high computational load and leaves behind unclean data. As a result, there is a lack of appropriate training datasets for text-to-video models. To address this problem, we present VidGen-1M, a superior training dataset for text-to-video models. Produced through a coarse-to-fine curation strategy, this dataset guarantees high-quality videos and detailed captions with excellent temporal consistency. When used to train the video generation model, this dataset has led to experimental results that surpass those obtained with other models.

📄 PDF Abstract BibTeX arXiv:2408.02629

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video GenerationVideo Generation

Similar Papers 제목 키워드 기반

Beyond Text Conditioning: A Systematic Study of MLLM-DiT Fusion for Video Generation

2026-08-14 · Yanbo Ding, Yijia Fan, Caihua Shan, Yifan Yang 외 arxiv

Diffusion Transformers (DiTs) have become the dominant paradigm for high-fidelity video generation, yet their ability to perform high-level semantic planning remains limited. While hybrid architectures integrating MLLMs …

Text-to-Video Generation

PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation

2025-12-31 · Yuanhao Cai, Kunpeng Li, Menglin Jia, Jialiang Wang 외 arxiv

Recent advances in text-to-video (T2V) generation have achieved good visual quality, yet synthesizing videos that faithfully follow physical laws remains an open challenge. Existing methods mainly based on graphics or pr…

Text-to-Video Generation

Endora: Video Generation Models as Endoscopy Simulators

2024-03-17 · Chenxin Li, Hengyu Liu, Yifan Liu, Brandon Y. Feng 외

Generative models hold promise for revolutionizing medical education, robot-assisted surgery, and data augmentation for machine learning. Despite progress in generating 2D medical images, the complex domain of clinical v…

Data AugmentationVideo Generation

Reducing Target Group Bias in Hate Speech Detectors

2021-12-07 · Darsh J Shah, Sinong Wang, Han Fang, Hao Ma 외

The ubiquity of offensive and hateful content on online fora necessitates the need for automatic solutions that detect such content competently across target groups. In this paper we show that text classification models …

text-classificationText Classification

CNVid-3.5M: Build, Filter, and Pre-Train the Large-Scale Public Chinese Video-Text Dataset

2023-01-01 · CVPR 2023 1 · Tian Gan, Qing Wang, Xingning Dong, Xiangyuan Ren 외

Owing to well-designed large-scale video-text datasets, recent years have witnessed tremendous progress in video-text pre-training. However, existing large-scale video-text datasets are mostly English-only. Though th…