Classifying Syntactic Regularities for Hundreds of Languages
This paper presents a comparison of classification methods for linguistic typology for the purpose of expanding an extensive, but sparse language resource: the World Atlas of Language Structures (WALS) (Dryer and Haspelmath, 2013). We experimented with a variety of regression and nearest-neighbor methods for use in classification over a set of 325 languages and six syntactic rules drawn from WALS. To classify each rule, we consider the typological features of the other five rules; linguistic features extracted from a word-aligned Bible in each language; and genealogical features (genus and family) of each language. In general, we find that propagating the majority label among all languages of the same genus achieves the best accuracy in label pre- diction. Following this, a logistic regression model that combines typological and linguistic features offers the next best performance. Interestingly, this model actually outperforms the majority labels among all languages of the same family.
Code (0)
등록된 구현이 없습니다.
Tasks
General ClassificationregressionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Classifying Syntactic Errors in Learner Language
We present a method for classifying syntactic errors in learner language, namely errors whose correction alters the morphosyntactic structure of a sentence. The methodology builds on the established Universal Dependencie…
ClassificationGeneral ClassificationGrammatical Error CorrectionSentenceUniversal and Independent: Multilingual Probing Framework for Exhaustive Model Interpretation and Evaluation
Linguistic analysis of language models is one of the ways to explain and describe their reasoning, weaknesses, and limitations. In the probing part of the model interpretability research, studies concern individual langu…
Probing Language ModelsMLCPD: A Unified Multi-Language Code Parsing Dataset with Universal AST Schema
We introduce the MultiLang Code Parser Dataset (MLCPD), a large-scale, language-agnostic dataset unifying syntactic and structural representations of code across ten major programming languages. MLCPD contains over seven…
Representation LearningCapturing syntactico-semantic regularities among terms: An application of the FrameNet methodology to terminology
Terminological databases do not always provide detailed information on the linguistic behaviour of terms, although this is important for potential users such as translators or students. In this paper we describe a projec…
Universal Topological Regularities of Syntactic Structures: Decoupling Efficiency from Optimization
Human syntactic structures are usually represented as graphs. Much research has focused on the mapping between such graphs and linguistic sequences, but less attention has been paid to the shapes of the graphs themselves…