paper-with-me

Papers

ComicsPAP: understanding comic strips by picking the correct panel

2025-03-11 · Emanuele Vivoli, Artemis Llabrés, Mohamed Ali Souibgui, Marco Bertini, Ernest Valveny Llobet, Dimosthenis Karatzas

Large multimodal models (LMMs) have made impressive strides in image captioning, VQA, and video comprehension, yet they still struggle with the intricate temporal and spatial cues found in comics. To address this gap, we introduce ComicsPAP, a large-scale benchmark designed for comic strip understanding. Comprising over 100k samples and organized into 5 subtasks under a Pick-a-Panel framework, ComicsPAP demands models to identify the missing panel in a sequence. Our evaluations, conducted under both multi-image and single-image protocols, reveal that current state-of-the-art LMMs perform near chance on these tasks, underscoring significant limitations in capturing sequential and contextual dependencies. To close the gap, we adapted LMMs for comic strip understanding, obtaining better results on ComicsPAP than 10x bigger models, demonstrating that ComicsPAP offers a robust resource to drive future research in multimodal comic comprehension.

📄 PDF Abstract BibTeX arXiv:2503.08561

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Comics for Everyone: Generating Accessible Text Descriptions for Comic Strips

2023-10-01 · Reshma Ramaprasad

Comic strips are a popular and expressive form of visual storytelling that can convey humor, emotion, and information. However, they are inaccessible to the BLV (Blind or Low Vision) community, who cannot perceive the im…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+1

Towards Faithful Reasoning in Comics for Small MLLMs

2026-01-06 · Chengcheng Feng, Haojie Yin, Yucheng Jin, Kaizhu Huang arxiv

Comic understanding presents a significant challenge for Multimodal Large Language Models (MLLMs), as the intended meaning of a comic often emerges from the joint interpretation of visual, textual, and social cues. This …

Visual Reasoning

One missing piece in Vision and Language: A Survey on Comics Understanding

2024-09-14 · Emanuele Vivoli, Mohamed Ali Souibgui, Andrey Barsky, Artemis Llabrés 외

Vision-language models have recently evolved into versatile systems capable of high performance across a range of tasks, such as document understanding, visual question answering, and grounding, often in zero-shot settin…

document understandingimage-classificationImage ClassificationInstance Segmentation+6

Semantic Similarity is a Spurious Measure of Comic Understanding: Lessons Learned from Hallucinations in a Benchmarking Experiment

2026-03-02 · Christopher Driggers-Ellis, Nachiketh Tibrewal, Rohit Bogulla, Harsh Khanna 외 arxiv

A system that enables blind or visually impaired users to access comics/manga would introduce a new medium of storytelling to this community. However, no such system currently exists. Generative vision-language models (V…

Semantic Similarity

Gaze Heads: How VLMs Look at What They Describe

2026-06-12 · Rohit Gandikota, David Bau arxiv

How a vision-language model internally solves the task of describing an image is far from obvious. We find that the model develops a specific mechanism for this: a small set of attention heads in its language-model backb…

Continuous Control