paper-with-me

Papers

SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling

2024-02-01 · Eileen Wang, Soyeon Caren Han, Josiah Poon

Visual storytelling aims to automatically generate a coherent story based on a given image sequence. Unlike tasks like image captioning, visual stories should contain factual descriptions, worldviews, and human social commonsense to put disjointed elements together to form a coherent and engaging human-writeable story. However, most models mainly focus on applying factual information and using taxonomic/lexical external knowledge when attempting to create stories. This paper introduces SCO-VIST, a framework representing the image sequence as a graph with objects and relations that includes human action motivation and its social interaction commonsense knowledge. SCO-VIST then takes this graph representing plot points and creates bridges between plot points with semantic and occurrence-based edge weights. This weighted story graph produces the storyline in a sequence of events using Floyd-Warshall's algorithm. Our proposed framework produces stories superior across multiple metrics in terms of visual grounding, coherence, diversity, and humanness, per both automatic and human evaluations.

📄 PDF Abstract BibTeX arXiv:2402.00319

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityImage CaptioningVisual GroundingVisual Storytelling

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

VISTA: A Controllable Platform for Generating and Auditing Egocentric Assistance Scenarios

2026-05-11 · Yu-Hsiang Liu, Yu-Chien Tang, An-Zi Yen arxiv

Evaluating whether AI agents can proactively assist humans in daily activities, ranging from routine household tasks to urgent safety-critical situations, requires diverse visual data. However, collecting such scenarios …

SocialIQA: Commonsense Reasoning about Social Interactions

2019-04-22 · Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras 외

We introduce Social IQa, the first largescale benchmark for commonsense reasoning about social situations. Social IQa contains 38,000 multiple choice questions for probing emotional and social intelligence in a variety o…

Common Sense ReasoningCoreference ResolutionMultiple-choiceQuestion Answering+1

Social IQa: Commonsense Reasoning about Social Interactions

2019-11-01 · IJCNLP 2019 11 · Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras 외

We introduce Social IQa, the first large-scale benchmark for commonsense reasoning about social situations. Social IQa contains 38,000 multiple choice questions for probing emotional and social intelligence in a variety …

Multiple-choiceQuestion AnsweringTransfer Learning

Imagine, Reason and Write: Visual Storytelling with Graph Knowledge and Relational Reasoning

2021-05-18 · The Thirty-Fifth AAAI Conference on Artificial Intelligence 2021 5 · Chunpu Xu, Min Yang, Chengming Li, Ying Shen 외

Visual storytelling is a task of creating a short story based on photo streams. Different from visual captions, stories contain not only factual descriptions, but also imaginary concepts that do not appear in the images.…

DiversityInformativenessRelational ReasoningVisual Storytelling

COSMO: Conditional SEQ2SEQ-based Mixture Model for Zero-Shot Commonsense Question Answering

2020-11-02 · COLING 2020 8 · Farhad Moghimifar, Lizhen Qu, Yue Zhuo, Mahsa Baktashmotlagh 외

Commonsense reasoning refers to the ability of evaluating a social situation and acting accordingly. Identification of the implicit causes and effects of a social context is the driving capability which can enable machin…

Question Answering