paper-with-me

Papers

Structured Graph Representations for Visual Narrative Reasoning: A Hierarchical Framework for Comics

2025-04-14 · Yi-Chun Chen

This paper presents a hierarchical knowledge graph framework for the structured understanding of visual narratives, focusing on multimodal media such as comics. The proposed method decomposes narrative content into multiple levels, from macro-level story arcs to fine-grained event segments. It represents them through integrated knowledge graphs that capture semantic, spatial, and temporal relationships. At the panel level, we construct multimodal graphs that link visual elements such as characters, objects, and actions with corresponding textual components, including dialogue and captions. These graphs are integrated across narrative levels to support reasoning over story structure, character continuity, and event progression. We apply our approach to a manually annotated subset of the Manga109 dataset and demonstrate its ability to support symbolic reasoning across diverse narrative tasks, including action retrieval, dialogue tracing, character appearance mapping, and panel timeline reconstruction. Evaluation results show high precision and recall across tasks, validating the coherence and interpretability of the framework. This work contributes a scalable foundation for narrative-based content analysis, interactive storytelling, and multimodal reasoning in visual media.

📄 PDF Abstract BibTeX arXiv:2506.10008

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge GraphsMultimodal Reasoning

Similar Papers 제목 키워드 기반

Robust Symbolic Reasoning for Visual Narratives via Hierarchical and Semantically Normalized Knowledge Graphs

2025-08-20 · Yi-Chun Chen arxiv

Understanding visual narratives such as comics requires structured representations that capture events, characters, and their relations across multiple levels of story organization. However, symbolic narrative graphs oft…

Knowledge Graphs

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding

2026-01-03 · Hyeonjeong Ha, Jinjin Ge, Bo Feng, Kaixin Ma 외 arxiv

Multimodal large language models (MLLMs) have achieved impressive progress in vision-language reasoning, yet their ability to understand temporally unfolding narratives in videos remains underexplored. True narrative und…

Self-Reasoning Agentic Framework for Narrative Product Grid-Collage Generation

2026-04-18 · Minyan Luo, Yuxin Zhang, Yifei Li, Xincan Wang 외 arxiv

Narrative-driven product photography has become a prevalent paradigm in modern marketing, as coherent visual storytelling helps convey product value and establishes emotional engagement with consumers. However, existing …

Visual StorytellingImage Generation

GreaseLM: Graph REASoning Enhanced Language Models

2021-09-29 · ICLR 2022 4 · Xikun Zhang, Antoine Bosselut, Michihiro Yasunaga, Hongyu Ren 외

Answering complex questions about textual narratives requires reasoning over both stated context and the world knowledge that underlies it. However, pretrained language models (LM), the foundation of most modern QA syste…

Knowledge GraphsMedical Question AnsweringMedQANegation+3

Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs

2025-07-29 · Saeed Ghorbani arxiv

We introduce Aether Weaver, a novel, integrated framework for multimodal narrative co-generation that overcomes limitations of sequential text-to-visual pipelines. Our system concurrently synthesizes textual narratives, …