paper-with-me

홈 › Papers

Mind the Gap: Assessing Wiktionary's Crowd-Sourced Linguistic Knowledge on Morphological Gaps in Two Related Languages

2025-06-21 · Jonathan Sakunkoo, Annabella Sakunkoo

Morphological defectivity is an intriguing and understudied phenomenon in linguistics. Addressing defectivity, where expected inflectional forms are absent, is essential for improving the accuracy of NLP tools in morphologically rich languages. However, traditional linguistic resources often lack coverage of morphological gaps as such knowledge requires significant human expertise and effort to document and verify. For scarce linguistic phenomena in under-explored languages, Wikipedia and Wiktionary often serve as among the few accessible resources. Despite their extensive reach, their reliability has been a subject of controversy. This study customizes a novel neural morphological analyzer to annotate Latin and Italian corpora. Using the massive annotated data, crowd-sourced lists of defective verbs compiled from Wiktionary are validated computationally. Our results indicate that while Wiktionary provides a highly reliable account of Italian morphological gaps, 7% of Latin lemmata listed as defective show strong corpus evidence of being non-defective. This discrepancy highlights potential limitations of crowd-sourced wikis as definitive sources of linguistic knowledge, particularly for less-studied phenomena and languages, despite their value as resources for rare linguistic features. By providing scalable tools and methods for quality assurance of crowd-sourced data, this work advances computational morphology and expands linguistic knowledge of defectivity in non-English, morphologically rich languages.

📄 PDF Abstract BibTeX arXiv:2506.17603

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Assessing the Quality of an Italian Crowdsourced Idiom Corpus:the Dodiom Experiment

2022-06-01 · LREC 2022 6 · Giuseppina Morza, Raffaele Manna, Johanna Monti

This paper describes how idiom-related language resources, collected through a crowdsourcing experiment carried out by means of Dodiom, a Game-with-a-purpose, have been analysed by language experts. The paper focuses on …

Filtering Wiktionary Triangles by Linear Mbetween Distributed Word Models

2016-05-01 · LREC 2016 5 · M{\'a}rton Makrai

Word translations arise in dictionary-like organization as well as via machine learning from corpora. The former is exemplified by Wiktionary, a crowd-sourced dictionary with editions in many languages. {\'A}cs et al. (2…

Translation

Are BabyLMs Second Language Learners?

2024-10-28 · Lukas Edman, Lisa Bylinina, Faeze Ghorbanpour, Alexander Fraser

This paper describes a linguistically-motivated approach to the 2024 edition of the BabyLM Challenge (Warstadt et al. 2023). Rather than pursuing a first language learning (L1) paradigm, we approach the challenge from a …

Sentence

MucLex: A German Lexicon for Surface Realisation

2020-05-01 · LREC 2020 5 · Kira Klimt, Daniel Braun, Daniela Schneider, Florian Matthes

Language resources for languages other than English are often scarce. Rule-based surface realisers need elaborate lexica in order to be able to generate correct language, especially in languages like German, which includ…

Text Generation

LingoTurk: managing crowdsourced tasks for psycholinguistics

2016-06-01 · NAACL 2016 6 · Florian Pusse, Asad Sayeed, Vera Demberg
Machine Translation