paper-with-me

Papers

Visual Story Post-Editing

2019-06-05 · ACL 2019 7 · Ting-Yao Hsu, Chieh-Yang Huang, Yen-Chia Hsu, Ting-Hao 'Kenneth' Huang

We introduce the first dataset for human edits of machine-generated visual stories and explore how these collected edits may be used for the visual story post-editing task. The dataset, VIST-Edit, includes 14,905 human edited versions of 2,981 machine-generated visual stories. The stories were generated by two state-of-the-art visual storytelling models, each aligned to 5 human-edited versions. We establish baselines for the task, showing how a relatively small set of human edits can be leveraged to boost the performance of large visual storytelling models. We also discuss the weak correlation between automatic evaluation scores and human ratings, motivating the need for new automatic metrics.

📄 PDF Abstract BibTeX arXiv:1906.01764

Code (1)

tingyaohsu/VIST-Edit 공식 구현

Tasks

Visual Storytelling

Similar Papers 제목 키워드 기반

Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models

2025-09-04 · Kiymet Akdemir, Jing Shi, Kushal Kafle, Brian Price 외 arxiv

Text-to-image diffusion models have demonstrated significant capabilities to generate diverse and detailed visuals in various domains, and story visualization is emerging as a particularly promising application. However,…

Story VisualizationStory Generation

From Shots to Stories: LLM-Assisted Video Editing with Unified Language Representations

2025-05-18 · Yuzhi Li, Haojun Xu, Fang Tian

Large Language Models (LLMs) and Vision-Language Models (VLMs) have demonstrated remarkable reasoning and generalization capabilities in video understanding; however, their application in video editing remains largely un…

Video EditingVideo Understanding

StoryState: Agent-Based State Control for Consistent and Editable Storybooks

2026-02-01 · Ayushman Sarkar, Zhenyu Yu, Wei Tang, Chu Chen 외 arxiv

Large multimodal models have enabled one-click storybook generation, where users provide a short description and receive a multi-page illustrated story. However, the underlying story state, such as characters, world sett…

Text-to-Image Generation

Towards Data-Driven Automatic Video Editing

2019-07-17 · Sergey Podlesnyy

Automatic video editing involving at least the steps of selecting the most valuable footage from points of view of visual quality and the importance of action filmed; and cutting the footage into a brief and coherent vis…

Imitation LearningVideo Editing

PosEdiOn: Post-Editing Assessment in PythOn

2020-11-01 · EAMT 2020 11 · Antoni Oliver, Sergi Alvarez, Toni Badia

There is currently an extended use of post-editing of machine translation (PEMT) in the translation industry. This is due to the increase in the demand of translation and to the significant improvements in quality achiev…

Machine TranslationNMTSentenceTranslation