paper-with-me

홈 › Papers

URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors

2017-04-01 · EACL 2017 4 · Patrick Littell, David R. Mortensen, Ke Lin, Katherine Kairis, Carlisle Turner, Lori Levin

We introduce the URIEL knowledge base for massively multilingual NLP and the lang2vec utility, which provides information-rich vector identifications of languages drawn from typological, geographical, and phylogenetic databases and normalized to have straightforward and consistent formats, naming, and semantics. The goal of URIEL and lang2vec is to enable multilingual NLP, especially on less-resourced languages and make possible types of experiments (especially but not exclusively related to NLP tasks) that are otherwise difficult or impossible due to the sparsity and incommensurability of the data sources. lang2vec vectors have been shown to reduce perplexity in multilingual language modeling, when compared to one-hot language identification vectors.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationLanguage ModelingLanguage ModellingMultilingual NLP

Similar Papers 제목 키워드 기반

URIEL+: Enhancing Linguistic Inclusion and Usability in a Typological and Multilingual Knowledge Base

2024-09-27 · Aditya Khan, Mason Shipton, David Anugraha, Kaiyao Duan 외

URIEL is a knowledge base offering geographical, phylogenetic, and typological vector representations for 7970 languages. It includes distance measures between these vectors for 4005 languages, which are accessible via t…

A Reproducibility Study on Quantifying Language Similarity: The Impact of Missing Values in the URIEL Knowledge Base

2024-05-17 · Hasti Toossi, Guo Qing Huai, Jinyu Liu, Eric Khiu 외

In the pursuit of supporting more languages around the world, tools that characterize properties of languages play a key role in expanding the existing multilingual NLP research. In this study, we focus on a widely used …

Missing ValuesMultilingual NLP

Simple Additions, Substantial Gains: Expanding Scripts, Languages, and Lineage Coverage in URIEL+

2025-10-31 · Mason Shipton, York Hay Ng, Aditya Khan, Phuong Hanh Hoang 외 arxiv

The URIEL+ linguistic knowledge base supports multilingual research by encoding languages through geographic, genetic, and typological vectors. However, data sparsity (e.g. missing feature types, incomplete language entr…

Cross-Lingual Transfer

Less is More: The Effectiveness of Compact Typological Language Representations

2025-09-24 · York Hay Ng, Phuong Hanh Hoang, En-Shiun Annie Lee arxiv

Linguistic feature datasets such as URIEL+ are valuable for modelling cross-lingual relationships, but their high dimensionality and sparsity, especially for low-resource languages, limit the effectiveness of distance me…

LinguAlchemy: Fusing Typological and Geographical Elements for Unseen Language Generalization

2024-01-11 · Muhammad Farid Adilazuarda, Samuel Cahyawijaya, Alham Fikri Aji, Genta Indra Winata 외

Pretrained language models (PLMs) have become remarkably adept at task and language generalization. Nonetheless, they often fail when faced with unseen languages. In this work, we present LinguAlchemy, a regularization m…

intent-classificationIntent ClassificationLanguage ModellingNews Classification+1