paper-with-me

Papers

DuwatBench: Bridging Language and Visual Heritage through an Arabic Calligraphy Benchmark for Multimodal Understanding

2026-01-27 · Shubham Patle, Sara Ghaboura, Hania Tariq, Mohammad Usman Khan, Omkar Thawakar, Rao Muhammad Anwer, Salman Khan arxiv

Arabic calligraphy represents one of the richest visual traditions of the Arabic language, blending linguistic meaning with artistic form. Although multimodal models have advanced across languages, their ability to process Arabic script, especially in artistic and stylized calligraphic forms, remains largely unexplored. To address this gap, we present DuwatBench, a benchmark of 1,272 curated samples containing about 1,475 unique words across six classical and modern calligraphic styles, each paired with sentence-level detection annotations. The dataset reflects real-world challenges in Arabic writing, such as complex stroke patterns, dense ligatures, and stylistic variations that often challenge standard text recognition systems. Using DuwatBench, we evaluated 13 leading Arabic and multilingual multimodal models and showed that while they perform well on clean text, they struggle with calligraphic variation, artistic distortions, and precise visual-text alignment. By publicly releasing DuwatBench and its annotations, we aim to advance culturally grounded multimodal research, foster fair inclusion of the Arabic language and visual heritage in AI systems, and support continued progress in this area. Our dataset (https://huggingface.co/datasets/MBZUAI/DuwatBench) and evaluation suit (https://github.com/mbzuai-oryx/DuwatBench) are publicly available.

📄 PDF Abstract BibTeX arXiv:2601.19898

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution

2025-05-16 · Junyi Yuan, Jian Zhang, Fangyu Wu, Dongming Lu 외

China has a long and rich history, encompassing a vast cultural heritage that includes diverse multimodal information, such as silk patterns, Dunhuang murals, and their associated historical narratives. Cross-modal retri…

Cross-Modal RetrievalImage to textImage-to-Text RetrievalRetrieval+1

Diffusion Based Augmentation for Captioning and Retrieval in Cultural Heritage

2023-08-14 · Dario Cioni, Lorenzo Berlincioni, Federico Becattini, Alberto del Bimbo

Cultural heritage applications and advanced machine learning models are creating a fruitful synergy to provide effective and accessible ways of interacting with artworks. Smart audio-guides, personalized art-related cont…

Image CaptioningRetrieval

Quantum est in Libris: Navigating Archives with GenAI, Uncovering Tension Between Preservation and Innovation

2025-09-25 · Mar Canet Sola, Varvara Guljajeva arxiv

"Quantum est in libris" explores the intersection of the archaic and the modern. On one side, there are manuscript materials from the Estonian National Museum's (ERM) more than century-old archive describing the life exp…

ChinaHeritaQA: A Culturally-Grounded Visual Question Answering Dataset for World Heritage Sites in China

2026-06-08 · Yi Zhang, Bolei Ma, Yong Cao, Chengyan Wu 외 arxiv

We introduce ChinaHeritaQA, a multimodal benchmark dataset for evaluating the cultural reasoning abilities of vision-language models (VLMs) on UNESCO World Heritage sites in China. The dataset comprises 2,279 in-the-wild…

Visual Question Answering

LLM-Driven Completeness and Consistency Evaluation for Cultural Heritage Data Augmentation in Cross-Modal Retrieval

2025-11-09 · Jian Zhang, Junyi Guo, Junyi Yuan, Huanda Lu 외 arxiv

Cross-modal retrieval is essential for interpreting cultural heritage data, but its effectiveness is often limited by incomplete or inconsistent textual descriptions, caused by historical data loss and the high cost of e…

Cross-Modal RetrievalData Augmentation