paper-with-me

홈 › Papers

TRIDIS: A Comprehensive Medieval and Early Modern Corpus for HTR and NER

2025-03-25 · Sergio Torres Aguilar

This paper introduces TRIDIS (Tria Digita Scribunt), an open-source corpus of medieval and early modern manuscripts. TRIDIS aggregates multiple legacy collections (all published under open licenses) and incorporates large metadata descriptions. While prior publications referenced some portions of this corpus, here we provide a unified overview with a stronger focus on its constitution. We describe (i) the narrative, chronological, and editorial background of each major sub-corpus, (ii) its semi-diplomatic transcription rules (expansion, normalization, punctuation), (iii) a strategy for challenging out-of-domain test splits driven by outlier detection in a joint embedding space, and (iv) preliminary baseline experiments using TrOCR and MiniCPM2.5 comparing random and outlier-based test partitions. Overall, TRIDIS is designed to stimulate joint robust Handwritten Text Recognition (HTR) and Named Entity Recognition (NER) research across medieval and early modern textual heritage.

📄 PDF Abstract BibTeX arXiv:2503.22714

Code (0)

등록된 구현이 없습니다.

Tasks

Handwritten Text RecognitionHTRnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NEROutlier Detection

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Multi-Head Attention 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Residual Connection 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

People and Places of Historical Europe: Bootstrapping Annotation Pipeline and a New Corpus of Named Entities in Late Medieval Texts

2023-05-26 · Vít Novotný, Kristýna Luger, Michal Štefánik, Tereza Vrabcová 외

Although pre-trained named entity recognition (NER) models are highly accurate on modern corpora, they underperform on historical texts due to differences in language OCR errors. In this work, we develop a new NER corpus…

Information Retrievalnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+5

A Dataset for Metaphor Detection in Early Medieval Hebrew Poetry

2024-02-27 · Michael Toker, Oren Mishali, Ophir Münz-Manor, Benny Kimelfeld 외

There is a large volume of late antique and medieval Hebrew texts. They represent a crucial linguistic and cultural bridge between Biblical and modern Hebrew. Poetry is prominent in these texts and one of its main haract…

LiMe: a Latin Corpus of Late Medieval Criminal Sentences

2024-04-19 · Alessandra Bassani, Beatrice Del Bo, Alfio Ferrara, Marta Mangini 외

The Latin language has received attention from the computational linguistics research community, which has built, over the years, several valuable resources, ranging from detailed annotated corpora to sophisticated tools…

Language ModelingLanguage Modelling

Material Philology Meets Digital Onomastic Lexicography: The NordiCon Database of Medieval Nordic Personal Names in Continental Sources

2020-05-01 · LREC 2020 5 · Michelle Waldisp{\"u}hl, Dana Dannells, Lars Borin

We present NordiCon, a database containing medieval Nordic personal names attested in Continental sources. The database combines formally interpreted and richly interlinked onomastic data with digitized versions of the m…

Lemmatization

Nunc profana tractemus. Detecting Code-Switching in a Large Corpus of 16th Century Letters

2022-06-01 · LREC 2022 6 · Martin Volk, Lukas Fischer, Patricia Scheurer, Bernard Silvan Schroffenegger 외

This paper is based on a collection of 16th century letters from and to the Zurich reformer Heinrich Bullinger. Around 12,000 letters of this exchange have been preserved, out of which 3100 have been professionally edite…

Handwritten Text RecognitionMachine TranslationSentence