paper-with-me

홈 › Papers

Simple Additions, Substantial Gains: Expanding Scripts, Languages, and Lineage Coverage in URIEL+

2025-10-31 · Mason Shipton, York Hay Ng, Aditya Khan, Phuong Hanh Hoang, Xiang Lu, A. Seza Doğruöz, En-Shiun Annie Lee arxiv

The URIEL+ linguistic knowledge base supports multilingual research by encoding languages through geographic, genetic, and typological vectors. However, data sparsity (e.g. missing feature types, incomplete language entries, and limited genealogical coverage) remains prevalent. This limits the usefulness of URIEL+ in cross-lingual transfer, particularly for supporting low-resource languages. To address this sparsity, we extend URIEL+ by introducing script vectors to represent writing system properties for 7,488 languages, integrating Glottolog to add 18,710 additional languages, and expanding lineage imputation for 26,449 languages by propagating typological and script features across genealogies. These improvements reduce feature sparsity by 14% for script vectors, increase language coverage by up to 19,015 languages (1,007%), and boost imputation quality metrics by up to 35%. Our benchmark on cross-lingual transfer tasks (oriented around low-resource languages) shows occasionally divergent performance compared to URIEL+, with performance gains up to 6% in certain setups.

📄 PDF Abstract BibTeX arXiv:2510.27183

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Lingual Transfer

Similar Papers 제목 키워드 기반

MineProt: modern application for custom protein curation

2022-12-14 · Yunchi Zhu, Chengda Tong, Zuohan Zhao, Zuhong Lu

AI systems represented by AlphaFold are rapidly expanding the scale of protein structure modelling data, and the MineProt project provides an effective solution for custom curation of these novel high-throughput data. It…

On Romanization for Model Transfer Between Scripts in Neural Machine Translation

2020-09-30 · Findings of the Association for Computational Linguistics 2020 · Chantal Amrhein, Rico Sennrich

Transfer learning is a popular strategy to improve the quality of low-resource machine translation. For an optimal transfer of the embedding layer, the child and parent model should share a substantial part of the vocabu…

Machine TranslationTransfer LearningTranslation

Learning a performance metric of Buchberger's algorithm

2021-06-07 · Jelena Mojsilović, Dylan Peifer, Sonja Petrović

What can be (machine) learned about the complexity of Buchberger's algorithm? Given a system of polynomials, Buchberger's algorithm computes a Gr\"obner basis of the ideal these polynomials generate using an iterative pr…

regression

UNKs Everywhere: Adapting Multilingual Language Models to New Scripts

2020-12-31 · EMNLP 2021 11 · Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, Sebastian Ruder

Massively multilingual language models such as multilingual BERT offer state-of-the-art cross-lingual transfer performance on a range of NLP tasks. However, due to limited capacity and large differences in pretraining da…

Cross-Lingual Transfer

LLMs can generate robotic scripts from goal-oriented instructions in biological laboratory automation

2023-04-18 · Takashi Inagaki, Akari Kato, Koichi Takahashi, Haruka Ozaki 외

The use of laboratory automation by all researchers may substantially accelerate scientific activities by humans, including those in the life sciences. However, computer programs to operate robots should be written to im…