paper-with-me

Papers

DiffuVST: Narrating Fictional Scenes with Global-History-Guided Denoising Models

2023-12-12 · Shengguang Wu, Mei Yuan, Qi Su

Recent advances in image and video creation, especially AI-based image synthesis, have led to the production of numerous visual scenes that exhibit a high level of abstractness and diversity. Consequently, Visual Storytelling (VST), a task that involves generating meaningful and coherent narratives from a collection of images, has become even more challenging and is increasingly desired beyond real-world imagery. While existing VST techniques, which typically use autoregressive decoders, have made significant progress, they suffer from low inference speed and are not well-suited for synthetic scenes. To this end, we propose a novel diffusion-based system DiffuVST, which models the generation of a series of visual descriptions as a single conditional denoising process. The stochastic and non-autoregressive nature of DiffuVST at inference time allows it to generate highly diverse narratives more efficiently. In addition, DiffuVST features a unique design with bi-directional text history guidance and multimodal adapter modules, which effectively improve inter-sentence coherence and image-to-text fidelity. Extensive experiments on the story generation task covering four fictional visual-story datasets demonstrate the superiority of DiffuVST over traditional autoregressive models in terms of both text quality and inference speed.

📄 PDF Abstract BibTeX arXiv:2312.07066

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingDiversityImage GenerationImage to textSentenceStory GenerationVisual Storytelling

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
Adapter 설명 없음

Similar Papers 제목 키워드 기반

Meet Your Favorite Character: Open-domain Chatbot Mimicking Fictional Characters with only a Few Utterances

2021-11-16 · ACL ARR November 2021 11 · Anonymous

In this paper, we consider mimicking fictional characters as a promising direction for building engaging conversation models. To this end, we present a new practical task where only a few utterances of each fictional cha…

ChatbotRetrieval

Meet Your Favorite Character: Open-domain Chatbot Mimicking Fictional Characters with only a Few Utterances

2022-04-22 · NAACL 2022 7 · Seungju Han, Beomsu Kim, Jin Yong Yoo, Seokjun Seo 외

In this paper, we consider mimicking fictional characters as a promising direction for building engaging conversation models. To this end, we present a new practical task where only a few utterances of each fictional cha…

ChatbotRetrieval

Audiovisual speaker diarization of TV series

2018-12-18 · Xavier Bost, Georges Linarès, Serigne Gueye

Speaker diarization may be difficult to achieve when applied to narrative films, where speakers usually talk in adverse acoustic conditions: background music, sound effects, wide variations in intonation may hide the int…

speaker-diarizationSpeaker Diarization

Hierarchical memory decoder for visual narrating

2020-09-01 · IEEE Transactions on Circuits and Systems for Video Technology 2020 9 · Aming Wu, Yahong Han, Zhou Zhao, Yi Yang

Visual narrating focuses on generating semantic descriptions to summarize visual content of images or videos, e.g., visual captioning and visual storytelling. The challenge mainly lies in how to design a decoder to gener…

DecoderImage CaptioningVideo CaptioningVisual Storytelling

Movie101: A New Movie Understanding Benchmark

2023-05-20 · Zihao Yue, Qi Zhang, Anwen Hu, Liang Zhang 외

To help the visually impaired enjoy movies, automatic movie narrating systems are expected to narrate accurate, coherent, and role-aware plots when there are no speaking lines of actors. Existing works benchmark this cha…

Video Captioning