paper-with-me

홈 › Papers

URIEL+: Enhancing Linguistic Inclusion and Usability in a Typological and Multilingual Knowledge Base

2024-09-27 · Aditya Khan, Mason Shipton, David Anugraha, Kaiyao Duan, Phuong H. Hoang, Eric Khiu, A. Seza Doğruöz, En-Shiun Annie Lee

URIEL is a knowledge base offering geographical, phylogenetic, and typological vector representations for 7970 languages. It includes distance measures between these vectors for 4005 languages, which are accessible via the lang2vec tool. Despite being frequently cited, URIEL is limited in terms of linguistic inclusion and overall usability. To tackle these challenges, we introduce URIEL+, an enhanced version of URIEL and lang2vec that addresses these limitations. In addition to expanding typological feature coverage for 2898 languages, URIEL+ improves the user experience with robust, customizable distance calculations to better suit the needs of users. These upgrades also offer competitive performance on downstream tasks and provide distances that better align with linguistic distance studies.

📄 PDF Abstract BibTeX arXiv:2409.18472

Code (1)

Masonshipton25/URIELPlus 공식 구현

Methods 이 논문이 사용한 방법론

BASE 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Less is More: The Effectiveness of Compact Typological Language Representations

2025-09-24 · York Hay Ng, Phuong Hanh Hoang, En-Shiun Annie Lee arxiv

Linguistic feature datasets such as URIEL+ are valuable for modelling cross-lingual relationships, but their high dimensionality and sparsity, especially for low-resource languages, limit the effectiveness of distance me…

A Reproducibility Study on Quantifying Language Similarity: The Impact of Missing Values in the URIEL Knowledge Base

2024-05-17 · Hasti Toossi, Guo Qing Huai, Jinyu Liu, Eric Khiu 외

In the pursuit of supporting more languages around the world, tools that characterize properties of languages play a key role in expanding the existing multilingual NLP research. In this study, we focus on a widely used …

Missing ValuesMultilingual NLP

Simple Additions, Substantial Gains: Expanding Scripts, Languages, and Lineage Coverage in URIEL+

2025-10-31 · Mason Shipton, York Hay Ng, Aditya Khan, Phuong Hanh Hoang 외 arxiv

The URIEL+ linguistic knowledge base supports multilingual research by encoding languages through geographic, genetic, and typological vectors. However, data sparsity (e.g. missing feature types, incomplete language entr…

Cross-Lingual Transfer

Average Is Not Enough: Caveats of Multilingual Evaluation

2023-01-03 · Matúš Pikuliak, Marián Šimko

This position paper discusses the problem of multilingual evaluation. Using simple statistics, such as average language performance, might inject linguistic biases in favor of dominant language families into evaluation m…

Position

URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors

2017-04-01 · EACL 2017 4 · Patrick Littell, David R. Mortensen, Ke Lin, Katherine Kairis 외

We introduce the URIEL knowledge base for massively multilingual NLP and the lang2vec utility, which provides information-rich vector identifications of languages drawn from typological, geographical, and phylogenetic da…

Language IdentificationLanguage ModelingLanguage ModellingMultilingual NLP