paper-with-me

홈 › Papers

CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books

2025-07-14 · Marc Serra Ortega, Emanuele Vivoli, Artemis Llabrés, Dimosthenis Karatzas arxiv

This paper introduces CoSMo, a novel multimodal Transformer for Page Stream Segmentation (PSS) in comic books, a critical task for automated content understanding, as it is a necessary first stage for many downstream tasks like character analysis, story indexing, or metadata enrichment. We formalize PSS for this unique medium and curate a new 20,800-page annotated dataset. CoSMo, developed in vision-only and multimodal variants, consistently outperforms traditional baselines and significantly larger general-purpose vision-language models across F1-Macro, Panoptic Quality, and stream-level metrics. Our findings highlight the dominance of visual features for comic PSS macro-structure, yet demonstrate multimodal benefits in resolving challenging ambiguities. CoSMo establishes a new state-of-the-art, paving the way for scalable comic book analysis.

📄 PDF Abstract BibTeX arXiv:2507.10053

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Semantic Parsing of Interpage Relations

2022-05-26 · Mehmet Arif Demirtaş, Berke Oral, Mehmet Yasin Akpınar, Onur Deniz

Page-level analysis of documents has been a topic of interest in digitization efforts, and multimodal approaches have been applied to both classification and page stream segmentation. In this work, we focus on capturing …

ClassificationDependency ParsingPage Stream SegmentationSegmentation+1

Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models

2025-06-10 · Xuanchi Ren, Yifan Lu, Tianshi Cao, Ruiyuan Gao 외

Collecting and annotating real-world data for safety-critical physical AI systems, such as Autonomous Vehicle (AV), is time-consuming and costly. It is especially challenging to capture rare edge cases, which play a crit…

3D Lane Detection3D Object DetectionLane Detectionobject-detection+3

MMFace-DiT: A Dual-Stream Diffusion Transformer for High-Fidelity Multimodal Face Generation

2026-03-30 · Bharath Krishnamurthy, Ajita Rattani arxiv

Recent multimodal face generation models address the spatial control limitations of text-to-image diffusion models by augmenting text-based conditioning with spatial priors such as segmentation masks, sketches, or edge m…

Multi-modal Page Stream Segmentation with Convolutional Neural Networks

2019-09-27 · Lang Resources & Evaluation 2019 9 · Gregor Wiedemann, Gerhard Heyer

In recent years, (retro-)digitizing paper-based files became a major undertaking for private and public archives as well as an important task in electronic mailroom applications. As first steps, the workflow usually invo…

Optical Character RecognitionOptical Character Recognition (OCR)Page Stream SegmentationTransfer Learning

Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control

2025-03-18 · Nvidia, :, Hassan Abu Alhaija, Jose Alvarez 외

We introduce Cosmos-Transfer, a conditional world generation model that can generate world simulations based on multiple spatial control inputs of various modalities such as segmentation, depth, and edge. In the design, …