paper-with-me

홈 › Papers

Zooming into Comics: Region-Aware RL Improves Fine-Grained Comic Understanding in Vision-Language Models

2025-11-09 · Yule Chen, Yufan Ren, Sabine Süsstrunk arxiv

Complex visual narratives, such as comics, present a significant challenge to Vision-Language Models (VLMs). Despite excelling on natural images, VLMs often struggle with stylized line art, onomatopoeia, and densely packed multi-panel layouts. To address this gap, we introduce AI4VA-FG, the first fine-grained and comprehensive benchmark for VLM-based comic understanding. It spans tasks from foundational recognition and detection to high-level character reasoning and narrative construction, supported by dense annotations for characters, poses, and depth. Beyond that, we evaluate state-of-the-art proprietary models, including GPT-4o and Gemini-2.5, and open-source models such as Qwen2.5-VL, revealing substantial performance deficits across core tasks of our benchmarks and underscoring that comic understanding remains an unsolved challenge. To enhance VLMs' capabilities in this domain, we systematically investigate post-training strategies, including supervised fine-tuning on solutions (SFT-S), supervised fine-tuning on reasoning trajectories (SFT-R), and reinforcement learning (RL). Beyond that, inspired by the emerging "Thinking with Images" paradigm, we propose Region-Aware Reinforcement Learning (RARL) for VLMs, which trains models to dynamically attend to relevant regions through zoom-in operations. We observe that when applied to the Qwen2.5-VL model, RL and RARL yield significant gains in low-level entity recognition and high-level storyline ordering, paving the way for more accurate and efficient VLM applications in the comics domain.

📄 PDF Abstract BibTeX arXiv:2511.06490

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception

2026-02-12 · Lai Wei, Liangbo He, Jun Lan, Lingzhong Dong 외 arxiv

Multimodal Large Language Models (MLLMs) excel at broad visual understanding but still struggle with fine-grained perception, where decisive evidence is small and easily overwhelmed by global context. Recent "Thinking-wi…

Visual Reasoning

DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video Models

2025-12-01 · Patrick Kwon, Chen Chen arxiv

Current story visualization methods tend to position subjects solely by text and face challenges in maintaining artistic consistency. To address these limitations, we introduce DreamingComics, a layout-aware story visual…

Story Visualization

Real-time Surgical Environment Enhancement for Robot-Assisted Minimally Invasive Surgery Based on Super-Resolution

2020-11-08 · Ruoxi Wang, Dandan Zhang, QingBiao Li, Xiao-Yun Zhou 외

In Robot-Assisted Minimally Invasive Surgery (RAMIS), a camera assistant is normally required to control the position and zooming ratio of the laparoscope, following the surgeon's instructions. However, moving the laparo…

Depth EstimationGenerative Adversarial NetworkSuper-ResolutionVideo Super-Resolution

Painting Style-Aware Manga Colorization Based on Generative Adversarial Networks

2021-07-16 · Yugo Shimizu, Ryosuke Furuta, Delong Ouyang, Yukinobu Taniguchi 외

Japanese comics (called manga) are traditionally created in monochrome format. In recent years, in addition to monochrome comics, full color comics, a more attractive medium, have appeared. Unfortunately, color comics re…

Colorization

Emotion-Aware Speech Generation with Character-Specific Voices for Comics

2025-09-18 · Zhiwen Qian, Jinhua Liang, Huan Zhang arxiv

This paper presents an end-to-end pipeline for generating character-specific, emotion-aware speech from comics. The proposed system takes full comic volumes as input and produces speech aligned with each character's dial…