paper-with-me

Papers

FlashVideo:Flowing Fidelity to Detail for Efficient High-Resolution Video Generation

2025-02-07 · Shilong Zhang, Wenbo Li, Shoufa Chen, Chongjian Ge, Peize Sun, Yida Zhang, Yi Jiang, Zehuan Yuan, Binyue Peng, Ping Luo

DiT diffusion models have achieved great success in text-to-video generation, leveraging their scalability in model capacity and data scale. High content and motion fidelity aligned with text prompts, however, often require large model parameters and a substantial number of function evaluations (NFEs). Realistic and visually appealing details are typically reflected in high resolution outputs, further amplifying computational demands especially for single stage DiT models. To address these challenges, we propose a novel two stage framework, FlashVideo, which strategically allocates model capacity and NFEs across stages to balance generation fidelity and quality. In the first stage, prompt fidelity is prioritized through a low resolution generation process utilizing large parameters and sufficient NFEs to enhance computational efficiency. The second stage establishes flow matching between low and high resolutions, effectively generating fine details with minimal NFEs. Quantitative and visual results demonstrate that FlashVideo achieves state-of-the-art high resolution video generation with superior computational efficiency. Additionally, the two-stage design enables users to preview the initial output before committing to full resolution generation, thereby significantly reducing computational costs and wait times as well as enhancing commercial viability .

📄 PDF Abstract BibTeX arXiv:2502.05179

Code (1)

foundationvision/flashvideo 공식 구현 pytorch

Tasks

Computational EfficiencyText-to-Video GenerationVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

FlashVideo: A Framework for Swift Inference in Text-to-Video Generation

2023-12-30 · Bin Lei, Le Chen, Caiwen Ding

In the evolving field of machine learning, video generation has witnessed significant advancements with autoregressive-based transformer models and diffusion models, known for synthesizing dynamic and realistic scenes. H…

Text-to-Video GenerationVideo Generation

Pix2Streams: Dynamic Hydrology Maps from Satellite-LiDAR Fusion

2020-11-15 · Dolores Garcia, Gonzalo Mateo-Garcia, Hannes Bernhardt, Ron Hagensieker 외

Where are the Earth's streams flowing right now? Inland surface waters expand with floods and contract with droughts, so there is no one map of our streams. Current satellite approaches are limited to monthly observation…

Time Series Analysis

Spatial-Angular Representation Learning for High-Fidelity Continuous Super-Resolution in Diffusion MRI

2025-01-27 · Ruoyou Wu, Jian Cheng, Cheng Li, Juan Zou 외

Diffusion magnetic resonance imaging (dMRI) often suffers from low spatial and angular resolution due to inherent limitations in imaging hardware and system noise, adversely affecting the accurate estimation of microstru…

Diffusion MRIparameter estimationRepresentation LearningSuper-Resolution

Interpretable Detail-Fidelity Attention Network for Single Image Super-Resolution

2020-09-28 · Yuanfei Huang, Jie Li, Xinbo Gao, Yanting Hu 외

Benefiting from the strong capabilities of deep CNNs for feature representation and nonlinear mapping, deep-learning-based methods have achieved excellent performance in single image super-resolution. However, most exist…

Image Super-ResolutionSuper-Resolution

FiDeSR: High-Fidelity and Detail-Preserving One-Step Diffusion Super-Resolution

2026-03-03 · Aro Kim, Myeongjin Jang, Chaewon Moon, Youngjin Shin 외 arxiv

Diffusion-based approaches have recently driven remarkable progress in real-world image super-resolution (SR). However, existing methods still struggle to simultaneously preserve fine details and ensure high-fidelity rec…

Image Super-Resolution