paper-with-me

홈 › Papers

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation

2024-12-24 · Faraz Waseem, Muhammad Shahzad

An image may convey a thousand words, but a video composed of hundreds or thousands of image frames tells a more intricate story. Despite significant progress in multimodal large language models (MLLMs), generating extended videos remains a formidable challenge. As of this writing, OpenAI's Sora, the current state-of-the-art system, is still limited to producing videos that are up to one minute in length. This limitation stems from the complexity of long video generation, which requires more than generative AI techniques for approximating density functions essential aspects such as planning, story development, and maintaining spatial and temporal consistency present additional hurdles. Integrating generative AI with a divide-and-conquer approach could improve scalability for longer videos while offering greater control. In this survey, we examine the current landscape of long video generation, covering foundational techniques like GANs and diffusion models, video generation strategies, large-scale training datasets, quality metrics for evaluating long videos, and future research areas to address the limitations of the existing video generation capabilities. We believe it would serve as a comprehensive foundation, offering extensive information to guide future advancements and research in the field of long video generation.

📄 PDF Abstract BibTeX arXiv:2412.18688

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Semantic Video Segmentation : Exploring Inference Efficiency

2015-09-04 · Subarna Tripathi, Serge Belongie, Youngbae Hwang, Truong Nguyen

We explore the efficiency of the CRF inference beyond image level semantic segmentation and perform joint inference in video frames. The key idea is to combine best of two worlds: semantic co-labeling and more expressive…

Image SegmentationSegmentationSemantic SegmentationVideo Segmentation+1

Beyond Semantic Image Segmentation : Exploring Efficient Inference in Video

2015-07-01 · Subarna Tripathi, Serge Belongie, Truong Nguyen

We explore the efficiency of the CRF inference module beyond image level semantic segmentation. The key idea is to combine the best of two worlds of semantic co-labeling and exploiting more expressive models. Similar to …

Image SegmentationSegmentationSemantic SegmentationVideo Semantic Segmentation

A Video Is Not Worth a Thousand Words

2025-10-27 · Sam Pollard, Michael Wray arxiv

As we become increasingly dependent on vision language models (VLMs) to answer questions about the world around us, there is a significant amount of research devoted to increasing both the difficulty of video question an…

Video Question Answering

VESR-Net: The Winning Solution to Youku Video Enhancement and Super-Resolution Challenge

2020-03-04 · Jiale Chen, Xu Tan, Chaowei Shan, Sen Liu 외

This paper introduces VESR-Net, a method for video enhancement and super-resolution (VESR). We design a separate non-local module to explore the relations among video frames and fuse video frames efficiently, and a chann…

Super-ResolutionVideo Enhancement

Image and Information

2016-02-03 · Frank Nielsen

A well-known old adage says that {\em "A picture is worth a thousand words!"} (attributed to the Chinese philosopher Confucius ca 500 years BC). But more precisely, what do we mean by information in images? And how can i…