paper-with-me

홈 › Papers

StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation

2024-05-02 · Yupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng, Qibin Hou

For recent diffusion-based generative models, maintaining consistent content across a series of generated images, especially those containing subjects and complex details, presents a significant challenge. In this paper, we propose a new way of self-attention calculation, termed Consistent Self-Attention, that significantly boosts the consistency between the generated images and augments prevalent pretrained diffusion-based text-to-image models in a zero-shot manner. To extend our method to long-range video generation, we further introduce a novel semantic space temporal motion prediction module, named Semantic Motion Predictor. It is trained to estimate the motion conditions between two provided images in the semantic spaces. This module converts the generated sequence of images into videos with smooth transitions and consistent subjects that are significantly more stable than the modules based on latent spaces only, especially in the context of long video generation. By merging these two novel components, our framework, referred to as StoryDiffusion, can describe a text-based story with consistent images or videos encompassing a rich variety of contents. The proposed StoryDiffusion encompasses pioneering explorations in visual story generation with the presentation of images and videos, which we hope could inspire more research from the aspect of architectural modifications. Our code is made publicly available at https://github.com/HVision-NKU/StoryDiffusion.

📄 PDF Abstract BibTeX arXiv:2405.01434

Code (1)

hvision-nku/storydiffusion 공식 구현 pytorch

Tasks

motion predictionStory GenerationVideo Generation

Similar Papers 제목 키워드 기반

$π$-Attention: Online Efficient Sparse Transformers for Long-Context Modeling

2025-11-12 · Dong Liu, Yanxuan Yu arxiv

Sparse attention is crucial in long-context Transformers, which restricts each token to a limited neighborhood and thereby reduces the quadratic cost of full self-attention. Local windows capture nearby context effective…

Long-range modeling

Sketching as a Tool for Understanding and Accelerating Self-attention for Long Sequences

2021-12-10 · NAACL 2022 7 · Yifan Chen, Qi Zeng, Dilek Hakkani-Tur, Di Jin 외

Transformer-based models are not efficient in processing long sequences due to the quadratic space and time complexity of the self-attention modules. To address this limitation, Linformer and Informer are proposed to red…

Relational Attention Network for Crowd Counting

2019-10-01 · ICCV 2019 10 · Anran Zhang, Jiayi Shen, Zehao Xiao, Fan Zhu 외

Crowd counting is receiving rapidly growing research interests due to its potential application value in numerous real-world scenarios. However, due to various challenges such as occlusion, insufficient resolution and dy…

Crowd CountingDensity Estimation

You Only Sample (Almost) Once: Linear Cost Self-Attention Via Bernoulli Sampling

2021-11-18 · Zhanpeng Zeng, Yunyang Xiong, Sathya N. Ravi, Shailesh Acharya 외

Transformer-based models are widely used in natural language processing (NLP). Central to the transformer model is the self-attention mechanism, which captures the interactions of token pairs in the input sequences and d…

GPU

Cross-Modal Self-Attention Network for Referring Image Segmentation

2019-04-09 · CVPR 2019 6 · Linwei Ye, Mrigank Rochan, Zhi Liu, Yang Wang

We consider the problem of referring image segmentation. Given an input image and a natural language expression, the goal is to segment the object referred by the language expression in the image. Existing works in this …

Image SegmentationReferring ExpressionReferring Expression SegmentationReferring Video Object Segmentation+1