paper-with-me

홈 › Papers

bgGLUE: A Bulgarian General Language Understanding Evaluation Benchmark

2023-06-04 · Momchil Hardalov, Pepa Atanasova, Todor Mihaylov, Galia Angelova, Kiril Simov, Petya Osenova, Ves Stoyanov, Ivan Koychev, Preslav Nakov, Dragomir Radev

We present bgGLUE(Bulgarian General Language Understanding Evaluation), a benchmark for evaluating language models on Natural Language Understanding (NLU) tasks in Bulgarian. Our benchmark includes NLU tasks targeting a variety of NLP problems (e.g., natural language inference, fact-checking, named entity recognition, sentiment analysis, question answering, etc.) and machine learning tasks (sequence labeling, document-level classification, and regression). We run the first systematic evaluation of pre-trained language models for Bulgarian, comparing and contrasting results across the nine tasks in the benchmark. The evaluation results show strong performance on sequence labeling tasks, but there is a lot of room for improvement for tasks that require more complex reasoning. We make bgGLUE publicly available together with the fine-tuning and the evaluation code, as well as a public leaderboard at https://bgglue.github.io/, and we hope that it will enable further advancements in developing NLU models for Bulgarian.

📄 PDF Abstract BibTeX arXiv:2306.02349

Code (2)

bgGLUE/bgglue 공식 구현
bgglue/bgglue.github.io 공식 구현

Tasks

Fact Checkingnamed-entity-recognitionNamed Entity RecognitionNatural Language InferenceNatural Language UnderstandingQuestion AnsweringSentiment Analysis

Similar Papers 제목 키워드 기반

Bulgarian-English and English-Bulgarian Machine Translation: System Design and Evaluation

2017-09-01 · RANLP 2017 9 · Petya Osenova, Kiril Simov

The paper presents a deep factored machine translation (MT) system between English and Bulgarian languages in both directions. The MT system is hybrid. It consists of three main steps: (1) the source-language text is lin…

Machine TranslationTranslation

Detecting Multilingual COVID-19 Misinformation on Social Media via Contextualized Embeddings

2021-06-01 · NAACL (NLP4IF) 2021 6 · Subhadarshi Panda, Sarah Ita Levitan

We present machine learning classifiers to automatically identify COVID-19 misinformation on social media in three languages: English, Bulgarian, and Arabic. We compared 4 multitask learning models for this task and foun…

Misinformation

Comparative Analysis of Fine-tuned Deep Learning Language Models for ICD-10 Classification Task for Bulgarian Language

2021-09-01 · RANLP 2021 9 · Boris Velichkov, Sylvia Vassileva, Simeon Gerginov, Boris Kraychev 외

The task of automatic diagnosis encoding into standard medical classifications and ontologies, is of great importance in medicine - both to support the daily tasks of physicians in the preparation and reporting of clinic…

Evaluation of Off-the-Shelf Language Identification Tools on Bulgarian Social Media Posts

2022-09-01 · CLIB 2022 9 · Silvia Gargova, Irina Temnikova, Ivo Dzhumerov, Hristiana Nikolaeva

Automatic Language Identification (LI) is a widely addressed task, but not all users (for example linguists) have the means or interest to develop their own tool or to train the existing ones with their own data. There a…

Language Identification

Hear about Verbal Multiword Expressions in the Bulgarian and the Romanian Wordnets Straight from the Horse's Mouth

2019-08-01 · WS 2019 8 · Verginica Barbu Mititelu, Ivelina Stoyanova, Svetlozara Leseva, Maria Mitrofan 외

In this paper we focus on verbal multiword expressions (VMWEs) in Bulgarian and Romanian as reflected in the wordnets of the two languages. The annotation of VMWEs relies on the classification defined within the PARSEME …

ClassificationGeneral Classification