paper-with-me

Papers

Enhancing Human Pose Estimation in Ancient Vase Paintings via Perceptually-grounded Style Transfer Learning

2020-12-10 · Prathmesh Madhu, Angel Villar-Corrales, Ronak Kosti, Torsten Bendschus, Corinna Reinhardt, Peter Bell, Andreas Maier, Vincent Christlein

Human pose estimation (HPE) is a central part of understanding the visual narration and body movements of characters depicted in artwork collections, such as Greek vase paintings. Unfortunately, existing HPE methods do not generalise well across domains resulting in poorly recognized poses. Therefore, we propose a two step approach: (1) adapting a dataset of natural images of known person and pose annotations to the style of Greek vase paintings by means of image style-transfer. We introduce a perceptually-grounded style transfer training to enforce perceptual consistency. Then, we fine-tune the base model with this newly created dataset. We show that using style-transfer learning significantly improves the SOTA performance on unlabelled data by more than 6% mean average precision (mAP) as well as mean average recall (mAR). (2) To improve the already strong results further, we created a small dataset (ClassArch) consisting of ancient Greek vase paintings from the 6-5th century BCE with person and pose annotations. We show that fine-tuning on this data with a style-transferred model improves the performance further. In a thorough ablation study, we give a targeted analysis of the influence of style intensities, revealing that the model learns generic domain styles. Additionally, we provide a pose-based image retrieval to demonstrate the effectiveness of our method.

📄 PDF Abstract BibTeX arXiv:2012.05616

Code (1)

angelvillar96/STLPose 공식 구현 pytorch

Tasks

Image RetrievalPose EstimationRetrievalStyle TransferTransfer Learning

Similar Papers 제목 키워드 기반

VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery

2025-10-06 · Nonghai Zhang, Zeyu Zhang, Jiazi Wang, Yang Zhao 외 arxiv

Vision-Language Models (VLMs) have achieved significant progress in multimodal understanding tasks, demonstrating strong capabilities particularly in general tasks such as image captioning and visual reasoning. However, …

Visual Question AnsweringVisual ReasoningImage Captioning

VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery

2025-09-21 · Jinchao Ge, Tengfei Cheng, Biao Wu, Zeyu Zhang 외 arxiv

Understanding cultural heritage artifacts such as ancient Greek pottery requires expert-level reasoning that remains challenging for current MLLMs due to limited domain-specific data. We introduce VaseVQA, a benchmark of…

Visual Question AnsweringReinforcement Learning

VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery

2026-07-07 · Jiazi Wang, Nonghai Zhang, Qiushi Xie, Zeyu Zhang 외 arxiv

Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitization with natural-language artifact exploration. However, in cultural heritage domains such as ancient Gr…

Forensic Study of Paintings Through the Comparison of Fabrics

2025-06-25 · Juan José Murillo-Fuentes, Pablo M. Olmos, Laura Alba-Carcelén

The study of canvas fabrics in works of art is a crucial tool for authentication, attribution and conservation. Traditional methods are based on thread density map matching, which cannot be applied when canvases do not c…

Integrated Framework for Selecting and Enhancing Ancient Marathi Inscription Images from Stone, Metal Plate, and Paper Documents

2026-01-08 · Bapu D. Chendage, Rajivkumar S. Mente arxiv

Ancient script images often suffer from severe background noise, low contrast, and degradation caused by aging and environmental effects. In many cases, the foreground text and background exhibit similar visual character…

Image Enhancement