paper-with-me

Papers

Overlooked Data in Typological Databases: What Grambank Teaches Us About Gaps in Grammars

2022-06-01 · LREC 2022 6 · Jakob Lesage, Hannah J. Haynie, Hedvig Skirgård, Tobias Weber, Alena Witzlack-Makarevich

Typological databases can contain a wealth of information beyond the collection of linguistic properties across languages. This paper shows how information often overlooked in typological databases can inform the research community about the state of description of the world’s languages. We illustrate this using Grambank, a morphosyntactic typological database covering 2,467 language varieties and based on 3,951 grammatical descriptions. We classify and quantify the comments that accompany coded values in Grambank. We then aggregate these comments and the coded values to derive a level of description for 17 grammatical domains that Grambank covers (negation, adnominal modification, participant marking, tense, aspect, etc.). We show that the description level of grammatical domains varies across space and time. Information about gaps and uncertainties in the descriptive knowledge of grammatical domains within and across languages is essential for a correct analysis of data in typological databases and for the study of grammatical diversity more generally. When collected in a database, such information feeds into disciplines that focus on primary data collection, such as grammaticography and language documentation.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

DescriptiveDiversityNegation

Similar Papers 제목 키워드 기반

The Past, Present, and Future of Typological Databases in NLP

2023-10-20 · Emi Baylor, Esther Ploeger, Johannes Bjerva

Typological information has the potential to be beneficial in the development of NLP models, particularly for low-resource languages. Unfortunately, current large-scale typological databases, notably WALS and Grambank, a…

Language ModelingLanguage Modelling

Multilingual Gradient Word-Order Typology from Universal Dependencies

2024-02-02 · Emi Baylor, Esther Ploeger, Johannes Bjerva

While information from the field of linguistic typology has the potential to improve performance on NLP tasks, reliable typological data is a prerequisite. Existing typological databases, including WALS and Grambank, suf…

The taggedPBC: Annotating a massive parallel corpus for crosslinguistic investigations

2025-05-18 · Hiram Ring

Existing datasets available for crosslinguistic investigations have tended to focus on large amounts of data for a small group of languages or a small amount of data for a large number of languages. This means that claim…

POS

Does Typological Blinding Impede Cross-Lingual Sharing?

2021-01-28 · EACL 2021 2 · Johannes Bjerva, Isabelle Augenstein

Bridging the performance gap between high- and low-resource languages has been the focus of much previous work. Typological features from databases such as the World Atlas of Language Structures (WALS) are a prime candid…

Learning Language Representations for Typology Prediction

2017-07-29 · EMNLP 2017 9 · Chaitanya Malaviya, Graham Neubig, Patrick Littell

One central mystery of neural NLP is what neural models "know" about their subject matter. When a neural machine translation system learns to translate from one language to another, does it learn the syntax or semantics …

Machine TranslationNMTPredictionTranslation