paper-with-me

홈 › Papers

M2C: Towards Automatic Multimodal Manga Complement

2023-10-26 · Hongcheng Guo, Boyang Wang, Jiaqi Bai, Jiaheng Liu, Jian Yang, Zhoujun Li

Multimodal manga analysis focuses on enhancing manga understanding with visual and textual features, which has attracted considerable attention from both natural language processing and computer vision communities. Currently, most comics are hand-drawn and prone to problems such as missing pages, text contamination, and aging, resulting in missing comic text content and seriously hindering human comprehension. In other words, the Multimodal Manga Complement (M2C) task has not been investigated, which aims to handle the aforementioned issues by providing a shared semantic space for vision and language understanding. To this end, we first propose the Multimodal Manga Complement task by establishing a new M2C benchmark dataset covering two languages. First, we design a manga argumentation method called MCoT to mine event knowledge in comics with large language models. Then, an effective baseline FVP-M$^{2}$ using fine-grained visual prompts is proposed to support manga complement. Extensive experimental results show the effectiveness of FVP-M$^{2}$ method for Multimodal Mange Complement.

📄 PDF Abstract BibTeX arXiv:2310.17130

Code (1)

hc-guo/m2c 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Context-Informed Machine Translation of Manga using Multimodal Large Language Models

2024-11-04 · Philip Lippmann, Konrad Skublicki, Joshua Tanner, Shonosuke Ishiwatari 외

Due to the significant time and effort required for handcrafting translations, most manga never leave the domestic Japanese market. Automatic manga translation is a promising potential solution. However, it is a budding …

Machine TranslationTranslation

Towards Fully Automated Manga Translation

2020-12-28 · Ryota Hinami, Shonosuke Ishiwatari, Kazuhiko Yasuda, Yusuke Matsui

We tackle the problem of machine translation of manga, Japanese comics. Manga translation involves two important problems in machine translation: context-aware and multimodal translation. Since text and images are mixed …

Machine TranslationTranslation

MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding

2025-05-26 · Jeonghun Baek, Kazuki Egashira, Shota Onohara, Atsuyuki Miyai 외

Manga, or Japanese comics, is a richly multimodal narrative form that blends images and text in complex ways. Teaching large multimodal models (LMMs) to understand such narratives at a human-like level could help manga c…

Question AnsweringVisual Question Answering

Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding

2026-05-20 · Jeonghun Baek, Atsuyuki Miyai, Shota Onohara, Hikaru Ikuta 외 arxiv

Manga is a culturally distinctive multimodal medium and one of the most influential forms of Japanese popular culture. As AI systems increasingly target manga understanding, OCR, and translation, Manga109 has become a fo…

M2M-Gen: A Multimodal Framework for Automated Background Music Generation in Japanese Manga Using Large Language Models

2024-10-13 · Megha Sharma, Muhammad Taimoor Haseeb, Gus Xia, Yoshimasa Tsuruoka

This paper introduces M2M Gen, a multi modal framework for generating background music tailored to Japanese manga. The key challenges in this task are the lack of an available dataset or a baseline. To address these chal…

Emotion ClassificationMusic Generation