paper-with-me

Papers

Logion: Machine Learning for Greek Philology

2023-05-01 · Charlie Cowen-Breen, Creston Brooks, Johannes Haubold, Barbara Graziosi

This paper presents machine-learning methods to address various problems in Greek philology. After training a BERT model on the largest premodern Greek dataset used for this purpose to date, we identify and correct previously undetected errors made by scribes in the process of textual transmission, in what is, to our knowledge, the first successful identification of such errors via machine learning. Additionally, we demonstrate the model's capacity to fill gaps caused by material deterioration of premodern manuscripts and compare the model's performance to that of a domain expert. We find that best performance is achieved when the domain expert is provided with model suggestions for inspiration. With such human-computer collaborations in mind, we explore the model's interpretability and find that certain attention heads appear to encode select grammatical features of premodern Greek.

📄 PDF Abstract BibTeX arXiv:2305.01099

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Adam 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
WordPiece 설명 없음

Similar Papers 제목 키워드 기반

Graecia capta ferum victorem cepit. Detecting Latin Allusions to Ancient Greek Literature

2023-08-23 · Frederick Riemenschneider, Anette Frank

Intertextual allusions hold a pivotal role in Classical Philology, with Latin authors frequently referencing Ancient Greek texts. Until now, the automatic identification of these intertextual references has been constrai…

Sentence

Open Philology at the University of Leipzig

2014-05-01 · LREC 2014 5 · Frederik Baumgardt, Giuseppe Celano, Gregory R. Crane, Stella Dee 외

The Open Philology Project at the University of Leipzig aspires to re-assert the value of philology in its broadest sense. Philology signifies the widest possible use of the linguistic record to enable a deep understandi…

Optical Character Recognition (OCR)

Exploring Large Language Models for Classical Philology

2023-05-23 · Frederick Riemenschneider, Anette Frank

Recent advances in NLP have led to the creation of powerful language models for many languages including Ancient Greek and Latin. While prior work on Classical languages unanimously uses BERT, in this work we create four…

BenchmarkingDecoderLemmatization

Structural Similarity of Boundary Conditions and an Efficient Local Search Algorithm for Goal Conflict Identification

2021-02-23 · Hongzhen Zhong, Hai Wan, Weilin Luo, Zhanhao Xiao 외

In goal-oriented requirements engineering, goal conflict identification is of fundamental importance for requirements analysis. The task aims to find the feasible situations which make the goals diverge within the domain…

PENELOPIE: Enabling Open Information Extraction for the Greek Language through Machine Translation

2021-03-28 · EACL 2021 2 · Dimitris Papadopoulos, Nikolaos Papadakis, Nikolaos Matsatsinis

In this paper we present our submission for the EACL 2021 SRW; a methodology that aims at bridging the gap between high and low-resource languages in the context of Open Information Extraction, showcasing it on the Greek…

Machine TranslationNMTOpen Information ExtractionTranslation