Papers Visual Storytelling
“Visual Storytelling” 태그가 달린 논문 152편 · 필터 해제
Sidecar: Training-Free Semantic Reuse for Character-Consistent Free-form Visual Storytelling
Visual storytelling requires generating images that follow a narrative while preserving consistent character identities across frames. In free-form story generation, a character is fully described only when first introdu…
Visual StorytellingStory GenerationID-V2V: Identity-Preserving Video Restylization
In visual storytelling, human performances are central to creative intent and narrative meaning. However, preserving human identity and performance while enabling flexible visual edits remains challenging for generative …
Visual StorytellingFreeStory: Training-Free Character Consistency for Free-Form Visual Storytelling
Visual storytelling aims to generate image sequences that are both aligned with narrative prompts and consistent in character appearance across images. Recent training-free methods improve character consistency by reusin…
Visual StorytellingMangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation
End-to-end manga generation is a structured visual storytelling task that requires story decomposition, recurring character and scene grounding, page layout design, panel rendering, page composition, and lettering. Howev…
Visual StorytellingAttriStory: Fine-grained Attribute Realization for Visual Storytelling with Diffusion Models
Visual storytelling with diffusion models has made impressive strides in maintaining character consistency across narrative scenes. However, a critical gap remains: while these methods ensure a character remains consiste…
Visual StorytellingStory GenerationPuzzled By ChatGPT? No more! A Jigsaw Puzzle to Promote AI Literacy and Awareness
The rapid adoption of Generative AI, including LLM-based chatbots like ChatGPT, has highlighted the need for accessible ways to support public understanding and AI literacy. To address this need, we introduce a game-base…
Visual StorytellingSemantic-Structural Alignment for Generative Pictorial Charts
Traditional statistical graphics are precise but often lack the visual appeal, memorability, and engagement of pictorial charts. We present a generative framework for the automated synthesis of pictorial charts that brid…
Visual StorytellingImage EditingDreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior
Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with consistent characters, scenes, and transitions. However, existing a…
Visual StorytellingStory ContinuationSelf-Reasoning Agentic Framework for Narrative Product Grid-Collage Generation
Narrative-driven product photography has become a prevalent paradigm in modern marketing, as coherent visual storytelling helps convey product value and establishes emotional engagement with consumers. However, existing …
Visual StorytellingImage GenerationCANVAS: Continuity-Aware Narratives via Visual Agentic Storyboarding
Long-form visual storytelling requires maintaining continuity across shots, including consistent characters, stable environments, and smooth scene transitions. While existing generative models can produce strong individu…
Visual StorytellingExpressEdit: Fast Editing of Stylized Facial Expressions with Diffusion Models in Photoshop
Facial expressions of characters are a vital component of visual storytelling. While current AI image editing models hold promise for assisting artists in the task of stylized expression editing, these models introduce g…
Visual StorytellingImage EditingStoryBlender: Inter-Shot Consistent and Editable 3D Storyboard with Spatial-temporal Dynamics
Storyboarding is a core skill in visual storytelling for film, animation, and games. However, automating this process requires a system to achieve two properties that current approaches rarely satisfy simultaneously: int…
Visual StorytellingCustomized Visual Storytelling with Unified Multimodal LLMs
Multimodal story customization aims to generate coherent story flows conditioned on textual descriptions, reference identity images, and shot types. While recent progress in story generation has shown promising results, …
Visual StorytellingStory GenerationPersistent Story World Simulation with Continuous Character Customization
Story visualization has gained increasing attention in computer vision. However, current methods often fail to achieve a synergy between accurate character customization, semantic alignment, and continuous integration of…
Visual StorytellingStory VisualizationChArtist: Generating Pictorial Charts with Unified Spatial and Subject Control
A pictorial chart is an effective medium for visual storytelling, seamlessly integrating visual elements with data charts. However, creating such images is challenging because the flexibility of visual elements often con…
Visual StorytellingTowards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization
Unified vision-language models have made significant progress in multimodal understanding and generation, yet they largely fall short in producing multimodal interleaved outputs, which is a crucial capability for tasks l…
Text-to-Image GenerationReinforcement LearningVisual StorytellingVisual ReasoningMMCOMET: A Large-Scale Multimodal Commonsense Knowledge Graph for Contextual Reasoning
We present MMCOMET, the first multimodal commonsense knowledge graph (MMKG) that integrates physical, social, and eventive knowledge. MMCOMET extends the ATOMIC2020 knowledge graph to include a visual dimension, through …
Visual StorytellingImage CaptioningImage RetrievalStoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
Visual storytelling models that correctly ground entities in images may still hallucinate semantic relationships, generating incorrect dialogue attribution, character interactions, or emotional states. We introduce Story…
Visual StorytellingVisual GroundingIs Information Density Uniform when Utterances are Grounded on Perception and Discourse?
The Uniform Information Density (UID) hypothesis posits that speakers are subject to a communicative pressure to distribute information evenly within utterances, minimising surprisal variance. While this hypothesis has b…
Visual StorytellingAD-MIR: Bridging the Gap from Perception to Persuasion in Advertising Video Understanding via Structured Reasoning
Multimodal understanding of advertising videos is essential for interpreting the intricate relationship between visual storytelling and abstract persuasion strategies. However, despite excelling at general search, existi…
Visual StorytellingSemantic Retrieval