paper-with-me

Papers

Zero-Shot Captioning for Cultural Heritage: Automated Image Analysis of Traditional Indonesian Clothing

2026-06-11 · Anugrah Aidin Yotolembah, Novanto Yudistira, Gembong Edhi Setyawan arxiv

This paper presents Custom ZeroCLIP, a retrieval-augmented vision-language framework for zero-shot captioning of Indonesian traditional garments. The dataset contains 3,800 expert-annotated images from all 38 Indonesian provinces. Using a province-level inductive zero-shot protocol, the model is trained on 24 seen provinces, validated on 6 seen provinces, and evaluated on 8 unseen provinces. The framework combines a frozen CLIP ViT-B/32 image encoder, a CLIP text encoder, a BERT text encoder, and an LSTM caption decoder. During inference, unseen-province labels and captions are unavailable, and retrieval uses only captions from training provinces. No unseen-province image, label, or caption is used during training, validation, or retrieval-bank construction. Custom ZeroCLIP achieves a CLIPScore of 0.8536, BLEU-4 of 0.3342, and METEOR of 0.4859, outperforming existing baselines. Ablation results show that retrieval improves cultural vocabulary recovery with a 19.3\% METEOR gain, while human evaluation confirms stronger cultural accuracy and fluency. The results demonstrate the effectiveness of retrieval-augmented domain adaptation for culturally grounded caption generation in low-resource heritage settings. The dataset is publicly available at https://github.com/AnugrahAidinYotolembah/Traditional-Indonesian-Clothing-Captioning-Dataset.

📄 PDF Abstract BibTeX arXiv:2606.13275

Code (0)

등록된 구현이 없습니다.

Tasks

Domain Adaptation

Similar Papers 제목 키워드 기반

Automated Error Detection in Digitized Cultural Heritage Documents

2014-04-01 · WS 2014 4 · Kata G{\'a}bor, Beno{\^\i}t Sagot
Optical Character Recognition (OCR)Spelling Correction

LLM-Driven Completeness and Consistency Evaluation for Cultural Heritage Data Augmentation in Cross-Modal Retrieval

2025-11-09 · Jian Zhang, Junyi Guo, Junyi Yuan, Huanda Lu 외 arxiv

Cross-modal retrieval is essential for interpreting cultural heritage data, but its effectiveness is often limited by incomplete or inconsistent textual descriptions, caused by historical data loss and the high cost of e…

Cross-Modal RetrievalData Augmentation

Multimodal Cultural Heritage Knowledge Graph Extension with Language and Vision Models

2026-05-17 · Yang Zhang, Nada Mimouni, Jean-Claude Moissinac, Fayçal Hamdi arxiv

The preservation and interpretation of cultural heritage increasingly rely on digital technologies, among which Knowledge Graphs (KGs) stand out for their ability to structure vast amounts of data. However, the construct…

Knowledge Graph CompletionKnowledge Graphs

Is GPT-3 all you need for Visual Question Answering in Cultural Heritage?

2022-07-25 · Pietro Bongini, Federico Becattini, Alberto del Bimbo

The use of Deep Learning and Computer Vision in the Cultural Heritage domain is becoming highly relevant in the last few years with lots of applications about audio smart guides, interactive museums and augmented reality…

AllQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Automated Monitoring of Cultural Heritage Artifacts Using Semantic Segmentation

2025-11-25 · Andrea Ranieri, Giorgio Palmieri, Silvia Biasotti arxiv

This paper addresses the critical need for automated crack detection in the preservation of cultural heritage through semantic segmentation. We present a comparative study of U-Net architectures, using various convolutio…

Semantic SegmentationCrack Segmentation