paper-with-me

Papers Visual Storytelling

“Visual Storytelling” 태그가 달린 논문 152편 · 필터 해제

Sidecar: Training-Free Semantic Reuse for Character-Consistent Free-form Visual Storytelling

2026-08-27 · Sibo Dong, Sarah Adel Bargal arxiv

Visual storytelling requires generating images that follow a narrative while preserving consistent character identities across frames. In free-form story generation, a character is fully described only when first introdu…

Visual StorytellingStory Generation

ID-V2V: Identity-Preserving Video Restylization

2026-07-24 · Yuancheng Xu, Mingming He, Pablo Salamanca, Li Ma 외 hf

In visual storytelling, human performances are central to creative intent and narrative meaning. However, preserving human identity and performance while enabling flexible visual edits remains challenging for generative …

Visual Storytelling

FreeStory: Training-Free Character Consistency for Free-Form Visual Storytelling

2026-06-23 · Sibo Dong, Ismail Shaheen, Sarah Adel Bargal arxiv

Visual storytelling aims to generate image sequences that are both aligned with narrative prompts and consistent in character appearance across images. Recent training-free methods improve character consistency by reusin…

Visual Storytelling

MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation

2026-05-27 · Muyao Wang, Zeke Xie, Yanhao Chen, Lixin Xiu 외 arxiv

End-to-end manga generation is a structured visual storytelling task that requires story decomposition, recurring character and scene grounding, page layout design, panel rendering, page composition, and lettering. Howev…

Visual Storytelling

AttriStory: Fine-grained Attribute Realization for Visual Storytelling with Diffusion Models

2026-05-20 · Manogna Sreenivas, Rohit Kumar, Soma Biswas arxiv

Visual storytelling with diffusion models has made impressive strides in maintaining character consistency across narrative scenes. However, a critical gap remains: while these methods ensure a character remains consiste…

Visual StorytellingStory Generation

Puzzled By ChatGPT? No more! A Jigsaw Puzzle to Promote AI Literacy and Awareness

2026-05-19 · Francesca Padovani, Malvina Nissim arxiv

The rapid adoption of Generative AI, including LLM-based chatbots like ChatGPT, has highlighted the need for accessible ways to support public understanding and AI literacy. To address this need, we introduce a game-base…

Visual Storytelling

Semantic-Structural Alignment for Generative Pictorial Charts

2026-05-05 · Zhida Sun, Yulin Zhang, Zheng Gu, Min Lu 외 arxiv

Traditional statistical graphics are precise but often lack the visual appeal, memorability, and engagement of pictorial charts. We present a generative framework for the automated synthesis of pictorial charts that brid…

Visual StorytellingImage Editing

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior

2026-04-19 · Junjia Huang, Binbin Yang, Pengxiang Yan, Jiyang Liu 외 arxiv

Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with consistent characters, scenes, and transitions. However, existing a…

Visual StorytellingStory Continuation

Self-Reasoning Agentic Framework for Narrative Product Grid-Collage Generation

2026-04-18 · Minyan Luo, Yuxin Zhang, Yifei Li, Xincan Wang 외 arxiv

Narrative-driven product photography has become a prevalent paradigm in modern marketing, as coherent visual storytelling helps convey product value and establishes emotional engagement with consumers. However, existing …

Visual StorytellingImage Generation

CANVAS: Continuity-Aware Narratives via Visual Agentic Storyboarding

2026-04-15 · Ishani Mondal, Yiwen Song, Mihir Parmar, Palash Goyal 외 arxiv

Long-form visual storytelling requires maintaining continuity across shots, including consistent characters, stable environments, and smooth scene transitions. While existing generative models can produce strong individu…

Visual Storytelling

ExpressEdit: Fast Editing of Stylized Facial Expressions with Diffusion Models in Photoshop

2026-04-03 · Kenan Tang, Jiasheng Guo, Jeffrey Lin, Yao Qin arxiv

Facial expressions of characters are a vital component of visual storytelling. While current AI image editing models hold promise for assisting artists in the task of stylized expression editing, these models introduce g…

Visual StorytellingImage Editing

StoryBlender: Inter-Shot Consistent and Editable 3D Storyboard with Spatial-temporal Dynamics

2026-04-01 · Bingliang Li, Zhenhong Sun, Jiaming Bian, Yuehao Wu 외 arxiv

Storyboarding is a core skill in visual storytelling for film, animation, and games. However, automating this process requires a system to achieve two properties that current approaches rarely satisfy simultaneously: int…

Visual Storytelling

Customized Visual Storytelling with Unified Multimodal LLMs

2026-03-29 · Wei-Hua Li, Cheng Sun, Chu-Song Chen arxiv

Multimodal story customization aims to generate coherent story flows conditioned on textual descriptions, reference identity images, and shot types. While recent progress in story generation has shown promising results, …

Visual StorytellingStory Generation

Persistent Story World Simulation with Continuous Character Customization

2026-03-17 · Jinlu Zhang, Qiyun Wang, Baoxiang Du, Jiayi Ji 외 arxiv

Story visualization has gained increasing attention in computer vision. However, current methods often fail to achieve a synergy between accurate character customization, semantic alignment, and continuous integration of…

Visual StorytellingStory Visualization

ChArtist: Generating Pictorial Charts with Unified Spatial and Subject Control

2026-03-15 · Shishi Xiao, Tongyu Zhou, David Laidlaw, Gromit Yeuk-Yin Chan arxiv

A pictorial chart is an effective medium for visual storytelling, seamlessly integrating visual elements with data charts. However, creating such images is challenging because the flexibility of visual elements often con…

Visual Storytelling

Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization

2026-03-10 · Ming Nie, Chunwei Wang, Jianhua Han, Hang Xu 외 arxiv

Unified vision-language models have made significant progress in multimodal understanding and generation, yet they largely fall short in producing multimodal interleaved outputs, which is a crucial capability for tasks l…

Text-to-Image GenerationReinforcement LearningVisual StorytellingVisual Reasoning

MMCOMET: A Large-Scale Multimodal Commonsense Knowledge Graph for Contextual Reasoning

2026-03-01 · Eileen Wang, Hiba Arnaout, Dhita Pratama, Shuo Yang 외 arxiv

We present MMCOMET, the first multimodal commonsense knowledge graph (MMKG) that integrates physical, social, and eventive knowledge. MMCOMET extends the ATOMIC2020 knowledge graph to include a visual dimension, through …

Visual StorytellingImage CaptioningImage Retrieval

StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles

2026-02-25 · Daniel Oliveira, David Martins de Matos arxiv

Visual storytelling models that correctly ground entities in images may still hallucinate semantic relationships, generating incorrect dialogue attribution, character interactions, or emotional states. We introduce Story…

Visual StorytellingVisual Grounding

Is Information Density Uniform when Utterances are Grounded on Perception and Discourse?

2026-02-16 · Matteo Gay, Coleman Haley, Mario Giulianelli, Edoardo Ponti arxiv

The Uniform Information Density (UID) hypothesis posits that speakers are subject to a communicative pressure to distribute information evenly within utterances, minimising surprisal variance. While this hypothesis has b…

Visual Storytelling

AD-MIR: Bridging the Gap from Perception to Persuasion in Advertising Video Understanding via Structured Reasoning

2026-02-07 · Binxiao Xu, Junyu Feng, Xiaopeng Lin, Haodong Li 외 arxiv

Multimodal understanding of advertising videos is essential for interpreting the intricate relationship between visual storytelling and abstract persuasion strategies. However, despite excelling at general search, existi…

Visual StorytellingSemantic Retrieval
1–20 / 152 다음 →