paper-with-me

홈 › Papers

Sliced at SemEval-2022 Task 11: Bigger, Better? Massively Multilingual LMs for Multilingual Complex NER on an Academic GPU Budget

2022-07-01 · SemEval (NAACL) 2022 7 · Barbara Plank

Massively multilingual language models (MMLMs) have become a widely-used representation method, and multiple large MMLMs were proposed in recent years. A trend is to train MMLMs on larger text corpora or with more layers. In this paper we set out to test recent popular MMLMs on detecting semantically ambiguous and complex named entities with an academic GPU budget. Our submission of a single model for 11 languages on the SemEval Task 11 MultiCoNER shows that a vanilla transformer-CRF with XLM-R_{large} outperforms the more recent RemBERT, ranking 9th from 26 submissions in the multilingual track. Compared to RemBERT, the XLM-R model has the additional advantage to fit on a slice of a multi-instance GPU. As contrary to expectations and recent findings, we found RemBERT to not be the best MMLM, we further set out to investigate this discrepancy with additional experiments on multilingual Wikipedia NER data. While we expected RemBERT to have an edge on that dataset as it is closer to its pre-training data, surprisingly, our results show that this is not the case, suggesting that text domain match does not explain the discrepancy.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

GPUNERXLM-R

Methods 이 논문이 사용한 방법론

XLM-R XLM-R

Similar Papers 제목 키워드 기반

Smash at SemEval-2020 Task 7: Optimizing the Hyperparameters of ERNIE 2.0 for Humor Ranking and Rating

2020-12-01 · SEMEVAL 2020 · J. A. Meaney, Steven Wilson, Walid Magdy

The use of pre-trained language models such as BERT and ULMFiT has become increasingly popular in shared tasks, due to their powerful language modelling capabilities. Our entry to SemEval uses ERNIE 2.0, a language model…

ClassificationLanguage ModelingLanguage Modellingregression

Grunn2019 at SemEval-2019 Task 5: Shared Task on Multilingual Detection of Hate

2019-06-01 · SEMEVAL 2019 6 · Mike Zhang, Roy David, Leon Graumans, Gerben Timmerman

Hate speech occurs more often than ever and polarizes society. To help counter this polarization, SemEval 2019 organizes a shared task called the Multilingual Detection of Hate. The first task (A) is to decide whether a …

Fast Approximation of the Generalized Sliced-Wasserstein Distance

2022-10-19 · Dung Le, Huy Nguyen, Khai Nguyen, Trang Nguyen 외

Generalized sliced Wasserstein distance is a variant of sliced Wasserstein distance that exploits the power of non-linear projection through a given defining function to better capture the complex structures of the proba…

Babel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language Representations

2023-06-14 · Gregor Geigle, Radu Timofte, Goran Glavaš

Vision-and-language (VL) models with separate encoders for each modality (e.g., CLIP) have become the go-to models for zero-shot image classification and image-text retrieval. They are, however, mostly evaluated in Engli…

image-classificationImage ClassificationImage-text RetrievalMachine Translation+3

Bigger Buffer k-d Trees on Multi-Many-Core Systems

2015-12-09 · Fabian Gieseke, Cosmin Eugen Oancea, Ashish Mahabal, Christian Igel 외

A buffer k-d tree is a k-d tree variant for massively-parallel nearest neighbor search. While providing valuable speed-ups on modern many-core devices in case both a large number of reference and query points are given, …

Astronomy