paper-with-me

홈 › Papers

Predicting the Performance of Multilingual NLP Models

2021-10-17 · Anirudh Srinivasan, Sunayana Sitaram, Tanuja Ganu, Sandipan Dandapat, Kalika Bali, Monojit Choudhury

Recent advancements in NLP have given us models like mBERT and XLMR that can serve over 100 languages. The languages that these models are evaluated on, however, are very few in number, and it is unlikely that evaluation datasets will cover all the languages that these models support. Potential solutions to the costly problem of dataset creation are to translate datasets to new languages or use template-filling based techniques for creation. This paper proposes an alternate solution for evaluating a model across languages which make use of the existing performance scores of the model on languages that a particular task has test sets for. We train a predictor on these performance scores and use this predictor to predict the model's performance in different evaluation settings. Our results show that our method is effective in filling the gaps in the evaluation for an existing set of languages, but might require additional improvements if we want it to generalize to unseen languages.

📄 PDF Abstract BibTeX arXiv:2110.08875

Code (0)

등록된 구현이 없습니다.

Tasks

Multilingual NLP

Methods 이 논문이 사용한 방법론

Test 설명 없음
mBERT mBERT

Similar Papers 제목 키워드 기반

Unleashing the Multilingual Encoder Potential: Boosting Zero-Shot Performance via Probability Calibration

2023-10-08 · Ercong Nie, Helmut Schmid, Hinrich Schütze

Pretrained multilingual encoder models can directly perform zero-shot multilingual tasks or linguistic probing by reformulating the input examples into cloze-style prompts. This is accomplished by predicting the probabil…

Position

Multilingual and Multimodal Topic Modelling with Pretrained Embeddings

2022-11-15 · COLING 2022 10 · Elaine Zosa, Lidia Pivovarova

This paper presents M3L-Contrast -- a novel multimodal multilingual (M3L) neural topic model for comparable data that maps texts from multiple languages and images into a shared topic space. Our model is trained jointly …

Predicting Code-switching in Multilingual Communication for Immigrant Communities

2014-10-01 · WS 2014 10 · Evangelos Papalexakis, Dong Nguyen, A. Seza Do{\u{g}}ru{\"o}z
Language Identification

The Dabblers at SemEval-2018 Task 2: Multilingual Emoji Prediction

2018-06-01 · SEMEVAL 2018 6 · Larisa Alexa, Alina Loren{\textcommabelow{t}}, Daniela G{\^\i}fu, Tr 외

The {``}Multilingual Emoji Prediction{''} task focuses on the ability of predicting the correspondent emoji for a certain tweet. In this paper, we investigate the relation between words and emojis. In order to do that, w…

BIG-bench Machine LearningRelationTask 2

Unsupervised Translation Quality Estimation Exploiting Synthetic Data and Pre-trained Multilingual Encoder

2023-11-09 · Yuto Kuroda, Atsushi Fujita, Tomoyuki Kajiwara, Takashi Ninomiya

Translation quality estimation (TQE) is the task of predicting translation quality without reference translations. Due to the enormous cost of creating training data for TQE, only a few translation directions can benefit…

SentenceTranslation