paper-with-me

홈 › Papers

Toward accessible comics for blind and low vision readers

2024-07-11 · Christophe Rigaud, Jean-Christophe Burie, Samuel Petit

This work explores how to fine-tune large language models using prompt engineering techniques with contextual information for generating an accurate text description of the full story, ready to be forwarded to off-the-shelve speech synthesis tools. We propose to use existing computer vision and optical character recognition techniques to build a grounded context from the comic strip image content, such as panels, characters, text, reading order and the association of bubbles and characters. Then we infer character identification and generate comic book script with context-aware panel description including character's appearance, posture, mood, dialogues etc. We believe that such enriched content description can be easily used to produce audiobook and eBook with various voices for characters, captions and playing sound effects.

📄 PDF Abstract BibTeX arXiv:2407.08248

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Character RecognitionPrompt EngineeringSpeech Synthesis

Similar Papers 제목 키워드 기반

Comics for Everyone: Generating Accessible Text Descriptions for Comic Strips

2023-10-01 · Reshma Ramaprasad

Comic strips are a popular and expressive form of visual storytelling that can convey humor, emotion, and information. However, they are inaccessible to the BLV (Blind or Low Vision) community, who cannot perceive the im…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+1

From Panels to Prose: Generating Literary Narratives from Comics

2025-03-30 · Ragav Sachdeva, Andrew Zisserman

Comics have long been a popular form of storytelling, offering visually engaging narratives that captivate audiences worldwide. However, the visual nature of comics presents a significant barrier for visually impaired re…

Optical Character Recognition (OCR)

The Amazing Mysteries of the Gutter: Drawing Inferences Between Panels in Comic Book Narratives

2016-11-16 · CVPR 2017 7 · Mohit Iyyer, Varun Manjunatha, Anupam Guha, Yogarshi Vyas 외

Visual narrative is often a combination of explicit information and judicious omissions, relying on the viewer to supply missing details. In comics, most movements in time and space are hidden in the "gutters" between pa…

Semantic Similarity is a Spurious Measure of Comic Understanding: Lessons Learned from Hallucinations in a Benchmarking Experiment

2026-03-02 · Christopher Driggers-Ellis, Nachiketh Tibrewal, Rohit Bogulla, Harsh Khanna 외 arxiv

A system that enables blind or visually impaired users to access comics/manga would introduce a new medium of storytelling to this community. However, no such system currently exists. Generative vision-language models (V…

Semantic Similarity

Comics Datasets Framework: Mix of Comics datasets for detection benchmarking

2024-07-03 · Emanuele Vivoli, Irene Campaioli, Mariateresa Nardoni, Niccolò Biondi 외

Comics, as a medium, uniquely combine text and images in styles often distinct from real-world visuals. For the past three decades, computational research on comics has evolved from basic object detection to more sophist…

BenchmarkingObjectobject-detectionObject Detection+1