Toward accessible comics for blind and low vision readers
This work explores how to fine-tune large language models using prompt engineering techniques with contextual information for generating an accurate text description of the full story, ready to be forwarded to off-the-shelve speech synthesis tools. We propose to use existing computer vision and optical character recognition techniques to build a grounded context from the comic strip image content, such as panels, characters, text, reading order and the association of bubbles and characters. Then we infer character identification and generate comic book script with context-aware panel description including character's appearance, posture, mood, dialogues etc. We believe that such enriched content description can be easily used to produce audiobook and eBook with various voices for characters, captions and playing sound effects.
Code (0)
등록된 구현이 없습니다.
Tasks
Optical Character RecognitionPrompt EngineeringSpeech SynthesisSimilar Papers 제목 키워드 기반
Comics for Everyone: Generating Accessible Text Descriptions for Comic Strips
Comic strips are a popular and expressive form of visual storytelling that can convey humor, emotion, and information. However, they are inaccessible to the BLV (Blind or Low Vision) community, who cannot perceive the im…
Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+1From Panels to Prose: Generating Literary Narratives from Comics
Comics have long been a popular form of storytelling, offering visually engaging narratives that captivate audiences worldwide. However, the visual nature of comics presents a significant barrier for visually impaired re…
Optical Character Recognition (OCR)The Amazing Mysteries of the Gutter: Drawing Inferences Between Panels in Comic Book Narratives
Visual narrative is often a combination of explicit information and judicious omissions, relying on the viewer to supply missing details. In comics, most movements in time and space are hidden in the "gutters" between pa…
Semantic Similarity is a Spurious Measure of Comic Understanding: Lessons Learned from Hallucinations in a Benchmarking Experiment
A system that enables blind or visually impaired users to access comics/manga would introduce a new medium of storytelling to this community. However, no such system currently exists. Generative vision-language models (V…
Semantic SimilarityComics Datasets Framework: Mix of Comics datasets for detection benchmarking
Comics, as a medium, uniquely combine text and images in styles often distinct from real-world visuals. For the past three decades, computational research on comics has evolved from basic object detection to more sophist…
BenchmarkingObjectobject-detectionObject Detection+1