paper-with-me

Papers

From Phonology to Syntax: Unsupervised Linguistic Typology at Different Levels with Language Embeddings

2018-02-23 · NAACL 2018 6 · Johannes Bjerva, Isabelle Augenstein

A core part of linguistic typology is the classification of languages according to linguistic properties, such as those detailed in the World Atlas of Language Structure (WALS). Doing this manually is prohibitively time-consuming, which is in part evidenced by the fact that only 100 out of over 7,000 languages spoken in the world are fully covered in WALS. We learn distributed language representations, which can be used to predict typological properties on a massively multilingual scale. Additionally, quantitative and qualitative analyses of these language embeddings can tell us how language similarities are encoded in NLP models for tasks at different typological levels. The representations are learned in an unsupervised manner alongside tasks at three typological levels: phonology (grapheme-to-phoneme prediction, and phoneme reconstruction), morphology (morphological inflection), and syntax (part-of-speech tagging). We consider more than 800 languages and find significant differences in the language representations encoded, depending on the target task. For instance, although Norwegian Bokm{\aa}l and Danish are typologically close to one another, they are phonologically distant, which is reflected in their language embeddings growing relatively distant in a phonological task. We are also able to predict typological features in WALS with high accuracies, even for unseen language families.

📄 PDF Abstract BibTeX arXiv:1802.09375

Code (0)

등록된 구현이 없습니다.

Tasks

Morphological InflectionPart-Of-Speech Tagging

Similar Papers 제목 키워드 기반

Information flow, artificial phonology and typology

2021-02-01 · SCiL 2021 2 · Adamantios Gafos

A Portuguese-Spanish Corpus Annotated for Subject Realization and Referentiality

2012-05-01 · LREC 2012 5 · Luz Rello, Iria Gayo

This paper presents a comparable corpus of Portuguese and Spanish consisting of legal and health texts. We describe the annotation of zero subject, impersonal constructions and explicit subjects in the corpus. We annotat…

Coreference Resolution

Creating ConLangs to Probe the Metalinguistic Grammatical Knowledge of LLMs

2025-10-08 · Chihiro Taguchi, Richard Sproat arxiv

We present a system that uses LLMs as a tool in the development of Constructed Languages -- ConLangs, which we call IASC (Interactive Agentic System for ConLangs). The system is modular in that it creates each of the com…

KoBALT: Korean Benchmark For Advanced Linguistic Tasks

2025-05-22 · Hyopil Shin, Sangah Lee, Dongjun Jang, Wooseok Song 외

We introduce KoBALT (Korean Benchmark for Advanced Linguistic Tasks), a comprehensive linguistically-motivated benchmark comprising 700 multiple-choice questions spanning 24 phenomena across five linguistic domains: synt…

Multiple-choice

Divide and...conquer? On the limits of algorithmic approaches to syntactic semantic structure

2016-09-11 · Diego Gabriel Krivochen

In computer science, divide and conquer (D&C) is an algorithm design paradigm based on multi-branched recursion. A D&C algorithm works by recursively and monotonically breaking down a problem into sub problems of the sam…

valid