Less is More: The Effectiveness of Compact Typological Language Representations
Linguistic feature datasets such as URIEL+ are valuable for modelling cross-lingual relationships, but their high dimensionality and sparsity, especially for low-resource languages, limit the effectiveness of distance metrics. We propose a pipeline to optimize the URIEL+ typological feature space by combining feature selection and imputation, producing compact yet interpretable typological representations. We evaluate these feature subsets on linguistic distance alignment and downstream tasks, demonstrating that reduced-size representations of language typology can yield more informative distance metrics and improve performance in multilingual NLP applications.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
DivEMT: Neural Machine Translation Post-Editing Effort Across Typologically Diverse Languages
We introduce DivEMT, the first publicly available post-editing study of Neural Machine Translation (NMT) over a typologically diverse set of target languages. Using a strictly controlled setup, 18 professional translator…
Machine TranslationNMTTranslationThe interplay between language similarity and script on a novel multi-layer Algerian dialect corpus
Recent years have seen a rise in interest for cross-lingual transfer between languages with similar typology, and between languages of various scripts. However, the interplay between language similarity and difference in…
Cross-Lingual TransferPart-Of-Speech TaggingSentiment AnalysisTowards Typologically Aware Rescoring to Mitigate Unfaithfulness in Lower-Resource Languages
Multilingual large language models (LLMs) are known to more frequently generate non-faithful output in resource-constrained languages (Guerreiro et al., 2023 - arXiv:2303.16104), potentially because these typologically d…
Model SelectionA Principled Framework for Evaluating on Typologically Diverse Languages
Beyond individual languages, multilingual natural language processing (NLP) research increasingly aims to develop models that perform well across languages generally. However, evaluating these systems on all the world's …
Studying the Inductive Biases of RNNs with Synthetic Variations of Natural Languages
How do typological properties such as word order and morphological case marking affect the ability of neural sequence models to acquire the syntax of a language? Cross-linguistic comparisons of RNNs' syntactic performanc…
Object