paper-with-me

Papers

A Reproducibility Study on Quantifying Language Similarity: The Impact of Missing Values in the URIEL Knowledge Base

2024-05-17 · Hasti Toossi, Guo Qing Huai, Jinyu Liu, Eric Khiu, A. Seza Doğruöz, En-Shiun Annie Lee

In the pursuit of supporting more languages around the world, tools that characterize properties of languages play a key role in expanding the existing multilingual NLP research. In this study, we focus on a widely used typological knowledge base, URIEL, which aggregates linguistic information into numeric vectors. Specifically, we delve into the soundness and reproducibility of the approach taken by URIEL in quantifying language similarity. Our analysis reveals URIEL's ambiguity in calculating language distances and in handling missing values. Moreover, we find that URIEL does not provide any information about typological features for 31\% of the languages it represents, undermining the reliabilility of the database, particularly on low-resource languages. Our literature review suggests URIEL and lang2vec are used in papers on diverse NLP tasks, which motivates us to rigorously verify the database as the effectiveness of these works depends on the reliability of the information the tool provides.

📄 PDF Abstract BibTeX arXiv:2405.11125

Code (0)

등록된 구현이 없습니다.

Tasks

Missing ValuesMultilingual NLP

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

A reproducible experimental survey on biomedical sentence similarity: a string-based method sets the state of the art

2022-05-18 · Alicia Lara-Clares, Juan J. Lastra-Díaz, Ana Garcia-Serrano

This registered report introduces the largest, and for the first time, reproducible experimental survey on biomedical sentence similarity with the following aims: (1) to elucidate the state of the art of the problem; (2)…

Articlesnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+3

Quantifying perturbation impacts for large language models

2024-12-01 · Paulius Rauba, Qiyao Wei, Mihaela van der Schaar

We consider the problem of quantifying how an input perturbation impacts the outputs of large language models (LLMs), a fundamental task for model reliability and post-hoc interpretability. A key obstacle in this domain …

Semantic SimilaritySemantic Textual Similarity

Quantifying Reproducibility in NLP and ML

2021-09-02 · Anya Belz

Reproducibility has become an intensely debated topic in NLP and ML over recent years, but no commonly accepted way of assessing reproducibility, let alone quantifying it, has so far emerged. The assumption has been that…

Impacts of floating-point non-associativity on reproducibility for HPC and deep learning applications

2024-08-09 · Sanjif Shanmugavelu, Mathieu Taillefumier, Christopher Culver, Oscar Hernandez 외

Run to run variability in parallel programs caused by floating-point non-associativity has been known to significantly affect reproducibility in iterative algorithms, due to accumulating errors. Non-reproducibility can c…

Deep LearningGPUSensitivity

Foundations of LLM Knowledge Materialization: Termination, Reproducibility, Robustness

2025-10-08 · Luca Giordano, Simon Razniewski arxiv

Large Language Models (LLMs) encode substantial factual knowledge, yet measuring and systematizing this knowledge remains challenging. Converting it into structured format, for example through recursive extraction approa…

Semantic Similarity