paper-with-me

Papers

The GLAUx corpus: methodological issues in designing a long-term, diverse, multi-layered corpus of Ancient Greek

2021-08-01 · ACL (LChange) 2021 8 · Alek Keersmaekers

This paper describes the GLAUx project (“the Greek Language Automated”), an ongoing effort to develop a large long-term diachronic corpus of Greek, covering sixteen centuries of literary and non-literary material annotated with NLP methods. After providing an overview of related corpus projects and discussing the general architecture of the corpus, it zooms in on a number of larger methodological issues in the design of historical corpora. These include the encoding of textual variants, handling extralinguistic variation and annotating linguistic ambiguity. Finally, the long- and short-term perspectives of this project are discussed.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Probing the nature of an island constraint with a parsed corpus

2019-07-01 · LILT 2019 7 · Yusuke Kubota, Ai Kubota

This paper presents a case study of the use of the NINJAL Parsed Corpus of Modern Japanese (NPCMJ) for syntactic research. NPCMJ is the first phrase structure-based treebank for Japanese that is specifically designed for…

Why Chinese Web-as-Corpus is Wacky? Or: How Big Data is Killing Chinese Corpus Linguistics

2014-05-01 · LREC 2014 5 · Shu-Kai Hsieh

This paper aims to examine and evaluate the current development of using Web-as-Corpus (WaC) paradigm in Chinese corpus linguistics. I will argue that the unstable notion of wordhood in Chinese and the resulting diverse …

Chinese Word Segmentation

In Search of the Flocks: How to Perform Onomasiological Queries in an Ancient Greek Corpus?

2022-06-01 · LT4HALA (LREC) 2022 6 · Alek Keersmaekers, Toon Van Hal

This paper explores the possibilities of onomasiologically querying corpus data of Ancient Greek. The significance of the onomasiological approach has been highlighted in recent studies, yet the possibilities of performi…

Computational conceptual history of scientific concepts: From early digital methods to LLMs

2026-06-02 · Michael Zichert, Arno Simons arxiv

This article situates large language models (LLMs) within the longer history of computational approaches to concept analysis in the history, philosophy, and sociology of science (HPSS). We examine what LLMs add to existi…

Change Detection

SYN2015: Representative Corpus of Contemporary Written Czech

2016-05-01 · LREC 2016 5 · Michal K{\v{r}}en, V{\'a}clav Cvr{\v{c}}ek, Tom{\'a}{\v{s}} {\v{C}}apka, Anna {\v{C}}erm{\'a}kov{\'a} 외

The paper concentrates on the design, composition and annotation of SYN2015, a new 100-million representative corpus of contemporary written Czech. SYN2015 is a sequel of the representative corpora of the SYN series that…

text-classificationText Classification