paper-with-me

홈 › Papers

Transitional Adaptation of Pretrained Models for Visual Storytelling

2021-06-19 · CVPR 2021 1 · Youngjae Yu, Jiwan Chung, Heeseung Yun, Jongseok Kim, Gunhee Kim

Previous models for vision-to-language generation tasks usually pretrain a visual encoder and a language generator in the respective domains and jointly finetune them with the target task. However, this direct transfer practice may suffer from the discord between visual specificity and language fluency since they are often separately trained from large corpora of visual and text data with no common ground. In this work, we claim that a transitional adaptation task is required between pretraining and finetuning to harmonize the visual encoder and the language model for challenging downstream target tasks like visual storytelling. We propose a novel approach named Transitional Adaptation of Pretrained Model (TAPM) that adapts the multi-modal modules to each other with a simpler alignment task between visual inputs only with no need for text labels. Through extensive experiments, we show that the adaptation step significantly improves the performance of multiple language models for sequential video and image captioning tasks. We achieve new state-of-the-art performance on both language metrics and human evaluation in the multi-sentence description task of LSMDC 2019 and the image storytelling task of VIST. Our experiments reveal that this improvement in caption quality does not depend on the specific choice of language models.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningLanguage ModellingSentenceSpecificityText GenerationVisual Storytelling

Similar Papers 제목 키워드 기반

Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning

2024-08-12 · Yingjin Song, Denis Paperno, Albert Gatt

Visual storytelling systems generate multi-sentence stories from image sequences. In this task, capturing contextual information and bridging visual variation bring additional challenges. We propose a simple yet effectiv…

Contrastive LearningInformativenessSentenceVisual Storytelling

Visual Storytelling with Question-Answer Plans

2023-10-08 · Danyang Liu, Mirella Lapata, Frank Keller

Visual storytelling aims to generate compelling narratives from image sequences. Existing models often focus on enhancing the representation of the image sequence, e.g., with external knowledge sources or advanced graph …

Visual Storytelling

The Art of Storytelling: Multi-Agent Generative AI for Dynamic Multimodal Narratives

2024-09-17 · Samee Arif, Taimoor Arif, Muhammad Saad Haroon, Aamina Jamal Khan 외

This paper introduces the concept of an education tool that utilizes Generative Artificial Intelligence (GenAI) to enhance storytelling for children. The system combines GenAI-driven narrative co-creation, text-to-speech…

text-to-speechText to SpeechText-to-Video GenerationVideo Generation

Asymmetric Adaptation-based Real-time Fault Diagnosis Under Transitional Operating Conditions

2026-05-23 · Hongshuo Zhao, Zeyi Liu, Xiao He arxiv

Data streams in real-world industrial scenarios often contain transitional operating conditions that are uncovered during offline training, leading to significant distribution shifts. To bridge the gap between static off…

Domain GeneralizationTest-time AdaptationFault Diagnosis

OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory

2025-12-08 · Zhaochong An, Menglin Jia, Haonan Qiu, Zijian Zhou 외 arxiv

Storytelling in real-world videos often unfolds through multiple shots -- discontinuous yet semantically connected clips that together convey a coherent narrative. However, existing multi-shot video generation (MSV) meth…

Video Generation