paper-with-me

Papers

Comics for Everyone: Generating Accessible Text Descriptions for Comic Strips

2023-10-01 · Reshma Ramaprasad

Comic strips are a popular and expressive form of visual storytelling that can convey humor, emotion, and information. However, they are inaccessible to the BLV (Blind or Low Vision) community, who cannot perceive the images, layouts, and text of comics. Our goal in this paper is to create natural language descriptions of comic strips that are accessible to the visually impaired community. Our method consists of two steps: first, we use computer vision techniques to extract information about the panels, characters, and text of the comic images; second, we use this information as additional context to prompt a multimodal large language model (MLLM) to produce the descriptions. We test our method on a collection of comics that have been annotated by human experts and measure its performance using both quantitative and qualitative metrics. The outcomes of our experiments are encouraging and promising.

📄 PDF Abstract BibTeX arXiv:2310.00698

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelVisual Storytelling

Similar Papers 제목 키워드 기반

The Manga Whisperer: Automatically Generating Transcriptions for Comics

2024-01-18 · CVPR 2024 1 · Ragav Sachdeva, Andrew Zisserman

In the past few decades, Japanese comics, commonly referred to as Manga, have transcended both cultural and linguistic boundaries to become a true worldwide sensation. Yet, the inherent reliance on visual cues and illust…

ComicGAN: Text-to-Comic Generative Adversarial Network

2021-09-19 · Ben Proven-Bessel, Zilong Zhao, Lydia Chen

Drawing and annotating comic illustrations is a complex and difficult process. No existing machine learning algorithms have been developed to create comic illustrations based on descriptions of illustrations, or the dial…

Generative Adversarial NetworkImage Generation

Comics Datasets Framework: Mix of Comics datasets for detection benchmarking

2024-07-03 · Emanuele Vivoli, Irene Campaioli, Mariateresa Nardoni, Niccolò Biondi 외

Comics, as a medium, uniquely combine text and images in styles often distinct from real-world visuals. For the past three decades, computational research on comics has evolved from basic object detection to more sophist…

BenchmarkingObjectobject-detectionObject Detection+1

Toward accessible comics for blind and low vision readers

2024-07-11 · Christophe Rigaud, Jean-Christophe Burie, Samuel Petit

This work explores how to fine-tune large language models using prompt engineering techniques with contextual information for generating an accurate text description of the full story, ready to be forwarded to off-the-sh…

Optical Character RecognitionPrompt EngineeringSpeech Synthesis

Emotion-Aware Speech Generation with Character-Specific Voices for Comics

2025-09-18 · Zhiwen Qian, Jinhua Liang, Huan Zhang arxiv

This paper presents an end-to-end pipeline for generating character-specific, emotion-aware speech from comics. The proposed system takes full comic volumes as input and produces speech aligned with each character's dial…