paper-with-me

홈 › Papers

Visual Transformation Telling

2023-05-03 · Wanqing Cui, Xin Hong, Yanyan Lan, Liang Pang, Jiafeng Guo, Xueqi Cheng

Humans can naturally reason from superficial state differences (e.g. ground wetness) to transformations descriptions (e.g. raining) according to their life experience. In this paper, we propose a new visual reasoning task to test this transformation reasoning ability in real-world scenarios, called \textbf{V}isual \textbf{T}ransformation \textbf{T}elling (VTT). Given a series of states (i.e. images), VTT requires to describe the transformation occurring between every two adjacent states. Different from existing visual reasoning tasks that focus on surface state reasoning, the advantage of VTT is that it captures the underlying causes, e.g. actions or events, behind the differences among states. We collect a novel dataset to support the study of transformation reasoning from two existing instructional video datasets, CrossTask and COIN, comprising 13,547 samples. Each sample involves the key state images along with their transformation descriptions. Our dataset covers diverse real-world activities, providing a rich resource for training and evaluation. To construct an initial benchmark for VTT, we test several models, including traditional visual storytelling methods (CST, GLACNet, Densecap) and advanced multimodal large language models (LLaVA v1.5-7B, Qwen-VL-chat, Gemini Pro Vision, GPT-4o, and GPT-4). Experimental results reveal that even state-of-the-art models still face challenges in VTT, highlighting substantial areas for improvement.

📄 PDF Abstract BibTeX arXiv:2305.01928

Code (1)

hughplay/vtt 공식 구현 pytorch

Tasks

Dense Video CaptioningVideo CaptioningVisual ReasoningVisual Storytelling

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

The Art of Storytelling: Multi-Agent Generative AI for Dynamic Multimodal Narratives

2024-09-17 · Samee Arif, Taimoor Arif, Muhammad Saad Haroon, Aamina Jamal Khan 외

This paper introduces the concept of an education tool that utilizes Generative Artificial Intelligence (GenAI) to enhance storytelling for children. The system combines GenAI-driven narrative co-creation, text-to-speech…

text-to-speechText to SpeechText-to-Video GenerationVideo Generation

Vision Transformer Based Model for Describing a Set of Images as a Story

2022-10-06 · Zainy M. Malakan, Ghulam Mubashar Hassan, Ajmal Mian

Visual Story-Telling is the process of forming a multi-sentence story from a set of images. Appropriately including visual variation and contextual information captured inside the input images is one of the most challeng…

Language ModellingSentenceVisual Storytelling

World-State Transformations for Neuro-symbolic Interactive Storytelling

2026-05-23 · Santiago Góngora, Luis Chiruzzo, Gonzalo Méndez, Pablo Gervás arxiv

Large Language Models (LLMs) have changed the possibilities of Interactive Storytelling systems that process free-text user input. However, as more of these systems are built, evidence continues to mount regarding the st…

A Pipeline for Creative Visual Storytelling

2018-07-21 · WS 2018 6 · Stephanie M. Lukin, Reginald Hobbs, Clare R. Voss

Computational visual storytelling produces a textual description of events and interpretations depicted in a sequence of images. These texts are made possible by advances and cross-disciplinary approaches in natural lang…

Visual Storytelling

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?

2025-04-27 · Mohamed Gado, Towhid Taliee, Muhammad Memon, Dmitry Ignatov 외

Visual storytelling is an interdisciplinary field combining computer vision and natural language processing to generate cohesive narratives from sequences of images. This paper presents a novel approach that leverages re…

Visual GroundingVisual Storytelling