paper-with-me

Papers

Diffusion Based Augmentation for Captioning and Retrieval in Cultural Heritage

2023-08-14 · Dario Cioni, Lorenzo Berlincioni, Federico Becattini, Alberto del Bimbo

Cultural heritage applications and advanced machine learning models are creating a fruitful synergy to provide effective and accessible ways of interacting with artworks. Smart audio-guides, personalized art-related content and gamification approaches are just a few examples of how technology can be exploited to provide additional value to artists or exhibitions. Nonetheless, from a machine learning point of view, the amount of available artistic data is often not enough to train effective models. Off-the-shelf computer vision modules can still be exploited to some extent, yet a severe domain shift is present between art images and standard natural image datasets used to train such models. As a result, this can lead to degraded performance. This paper introduces a novel approach to address the challenges of limited annotated data and domain shifts in the cultural heritage domain. By leveraging generative vision-language models, we augment art datasets by generating diverse variations of artworks conditioned on their captions. This augmentation strategy enhances dataset diversity, bridging the gap between natural images and artworks, and improving the alignment of visual cues with knowledge from general-purpose datasets. The generated variations assist in training vision and language models with a deeper understanding of artistic characteristics and that are able to generate better captions with appropriate jargon.

📄 PDF Abstract BibTeX arXiv:2308.07151

Code (1)

ciodar/cultural-heritage-diffaug 공식 구현 pytorch

Tasks

Image CaptioningRetrieval

Similar Papers 제목 키워드 기반

Zero-Shot Captioning for Cultural Heritage: Automated Image Analysis of Traditional Indonesian Clothing

2026-06-11 · Anugrah Aidin Yotolembah, Novanto Yudistira, Gembong Edhi Setyawan arxiv

This paper presents Custom ZeroCLIP, a retrieval-augmented vision-language framework for zero-shot captioning of Indonesian traditional garments. The dataset contains 3,800 expert-annotated images from all 38 Indonesian …

Domain Adaptation

LLM-Driven Completeness and Consistency Evaluation for Cultural Heritage Data Augmentation in Cross-Modal Retrieval

2025-11-09 · Jian Zhang, Junyi Guo, Junyi Yuan, Huanda Lu 외 arxiv

Cross-modal retrieval is essential for interpreting cultural heritage data, but its effectiveness is often limited by incomplete or inconsistent textual descriptions, caused by historical data loss and the high cost of e…

Cross-Modal RetrievalData Augmentation

Cultural Heritage 3D Reconstruction with Diffusion Networks

2024-10-14 · Pablo Jaramillo, Ivan Sipiran

This article explores the use of recent generative AI algorithms for repairing cultural heritage objects, leveraging a conditional diffusion model designed to reconstruct 3D point clouds effectively. Our study evaluates …

3D ReconstructionDiversitySensitivity

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution

2025-05-16 · Junyi Yuan, Jian Zhang, Fangyu Wu, Dongming Lu 외

China has a long and rich history, encompassing a vast cultural heritage that includes diverse multimodal information, such as silk patterns, Dunhuang murals, and their associated historical narratives. Cross-modal retri…

Cross-Modal RetrievalImage to textImage-to-Text RetrievalRetrieval+1

Retrieval-Augmented Long-Context Translation for Cultural Image Captioning: Gators submission for AmericasNLP 2026 shared task

2026-05-20 · Aashish Dhawan, Christopher Driggers-Ellis, Dzmitry Kasinets, Daisy Zhe Wang 외 arxiv

We present the University of Florida Gators submission to the AmericasNLP 2026 shared task on cultural image captioning for Indigenous languages. Our two-stage pipeline generates a Spanish intermediate caption with Qwen2…

Data AugmentationImage Captioning