paper-with-me

Papers

Learning from scarce information: using synthetic data to classify Roman fine ware pottery

2021-07-03 · Santos J. Núñez Jareño, Daniël P. van Helden, Evgeny M. Mirkes, Ivan Y. Tyukin, Penelope M. Allison

In this article we consider a version of the challenging problem of learning from datasets whose size is too limited to allow generalisation beyond the training set. To address the challenge we propose to use a transfer learning approach whereby the model is first trained on a synthetic dataset replicating features of the original objects. In this study the objects were smartphone photographs of near-complete Roman terra sigillata pottery vessels from the collection of the Museum of London. Taking the replicated features from published profile drawings of pottery forms allowed the integration of expert knowledge into the process through our synthetic data generator. After this first initial training the model was fine-tuned with data from photographs of real vessels. We show, through exhaustive experiments across several popular deep learning architectures, different test priors, and considering the impact of the photograph viewpoint and excessive damage to the vessels, that the proposed hybrid approach enables the creation of classifiers with appropriate generalisation performance. This performance is significantly better than that of classifiers trained exclusively on the original data which shows the promise of the approach to alleviate the fundamental issue of learning from small datasets.

📄 PDF Abstract BibTeX arXiv:2107.01401

Code (0)

등록된 구현이 없습니다.

Tasks

Transfer Learning

Similar Papers 제목 키워드 기반

Leaving No Stone Unturned When Identifying and Classifying Verbal Multiword Expressions in the Romanian Wordnet

2019-07-01 · GWC 2019 7 · Verginica Mititelu, Maria Mitrofan

We present here the enhancement of the Romanian wordnet with a new type of information, very useful in language processing, namely types of verbal multi-word expressions. All verb literals made of two or more words are a…

TF3-RO-50M: Training Compact Romanian Language Models from Scratch on Synthetic Moral Microfiction

2026-01-15 · Mihai Dan Nadas, Laura Diosan, Andreea Tomescu, Andrei Piscoran arxiv

Recent advances in synthetic data generation have shown that compact language models can be trained effectively when the underlying corpus is structurally controlled and linguistically coherent. However, for morphologica…

Synthetic Data GenerationKnowledge Distillation

Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering

2025-11-02 · Vlad Negoita, Mihai Masala, Traian Rebedea arxiv

Large Language Models (LLMs) have recently exploded in popularity, often matching or outperforming human abilities on many tasks. One of the key factors in training LLMs is the availability and curation of high-quality d…

Translation Asymmetry in LLMs as a Data Augmentation Factor: A Case Study for 6 Romansh Language Varieties

2026-03-26 · Jannis Vamvas, Ignacio Pérez Prat, Angela Heldstab, Dominic P. Fischer 외 arxiv

Recent strategies for low-resource machine translation rely on LLMs to generate synthetic data based on text in higher-resource languages. We revisit this idea for Romansh, a language with 6 distinct varieties. LLMs tend…

Machine TranslationData Augmentation

Ab Initio: Automatic Latin Proto-word Reconstruction

2018-08-01 · COLING 2018 8 · Alina Maria Ciobanu, Liviu P. Dinu

Proto-word reconstruction is central to the study of language evolution. It consists of recreating the words in an ancient language from its modern daughter languages. In this paper we investigate automatic word form rec…