paper-with-me

홈 › Papers

Efficiently Identifying Low-Quality Language Subsets in Multilingual Datasets: A Case Study on a Large-Scale Multilingual Audio Dataset

2024-10-05 · Farhan Samir, Emily P. Ahn, Shreya Prakash, Márton Soskuthy, Vered Shwartz, Jian Zhu

Curating datasets that span multiple languages is challenging. To make the collection more scalable, researchers often incorporate one or more imperfect classifiers in the process, like language identification models. These models, however, are prone to failure, resulting in some language subsets being unreliable for downstream tasks. We introduce a statistical test, the Preference Proportion Test, for identifying such unreliable subsets. By annotating only 20 samples for a language subset, we're able to identify systematic transcription errors for 10 language subsets in a recent large multilingual transcribed audio dataset, X-IPAPack (Zhu et al., 2024). We find that filtering this low-quality data out when training models for the downstream task of phonetic transcription brings substantial benefits, most notably a 25.7% relative improvement on transcribing recordings in out-of-distribution languages. Our method lays a path forward for systematic and reliable multilingual dataset auditing.

📄 PDF Abstract BibTeX arXiv:2410.04292

Code (0)

등록된 구현이 없습니다.

Tasks

Language Identification

Similar Papers 제목 키워드 기반

MultiLegalPile: A 689GB Multilingual Legal Corpus

2023-06-03 · Joel Niklaus, Veton Matoshi, Matthias Stürmer, Ilias Chalkidis 외

Large, high-quality datasets are crucial for training Large Language Models (LLMs). However, so far, there are few datasets available for specialized critical domains such as law and the available ones are often only for…

Rethinking what Matters: Effective and Robust Multilingual Realignment for Low-Resource Languages

2025-11-09 · Quang Phuoc Nguyen, David Anugraha, Felix Gaschi, Jun Bin Cheng 외 arxiv

Realignment is a promising strategy to improve cross-lingual transfer in multilingual language models. However, empirical results are mixed and often unreliable, particularly for typologically distant or low-resource lan…

Cross-Lingual Transfer

MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model Pretraining

2025-07-02 · Zhixun Chen, Ping Guo, Wenhan Han, Yifan Zhang 외 arxiv

Data quality is a critical driver of large language model performance, yet existing model-based selection methods focus almost exclusively on English. We introduce MuRating, a scalable framework that transfers high-quali…

Test Set Quality in Multilingual LLM Evaluation

2025-08-04 · Chalamalasetti Kranti, Gabriel Bernier-Colborne, Yvan Gauthier, Sowmya Vajjala arxiv

Several multilingual benchmark datasets have been developed in a semi-automatic manner in the recent past to measure progress and understand the state-of-the-art in the multilingual capabilities of Large Language Models.…

Towards Building an End-to-End Multilingual Automatic Lyrics Transcription Model

2024-06-25 · Jiawen Huang, Emmanouil Benetos

Multilingual automatic lyrics transcription (ALT) is a challenging task due to the limited availability of labelled data and the challenges introduced by singing, compared to multilingual automatic speech recognition. Al…

Automatic Lyrics TranscriptionAutomatic Speech Recognitionspeech-recognitionSpeech Recognition