paper-with-me

홈 › Papers

Evaluation of Off-the-Shelf Language Identification Tools on Bulgarian Social Media Posts

2022-09-01 · CLIB 2022 9 · Silvia Gargova, Irina Temnikova, Ivo Dzhumerov, Hristiana Nikolaeva

Automatic Language Identification (LI) is a widely addressed task, but not all users (for example linguists) have the means or interest to develop their own tool or to train the existing ones with their own data. There are several off-the-shelf LI tools, but for some languages, it is unclear which tool is the best for specific types of text. This article presents a comparison of the performance of several off-the-shelf language identification tools on Bulgarian social media data. The LI tools are tested on a multilingual Twitter dataset (composed of 2966 tweets) and an existing Bulgarian Twitter dataset on the topic of fake content detection of 3350 tweets. The article presents the manual annotation procedure of the first dataset, a dis- cussion of the decisions of the two annotators, and the results from testing the 7 off-the-shelf LI tools on both datasets. Our findings show that the tool, which is the easiest for users with no programming skills, achieves the highest F1-Score on Bulgarian social media data, while other tools have very useful functionalities for Bulgarian social media texts.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language Identification

Similar Papers 제목 키워드 기반

HeLI-OTS, Off-the-shelf Language Identifier for Text

2022-06-01 · LREC 2022 6 · Tommi Jauhiainen, Heidi Jauhiainen, Krister Lindén

This paper introduces HeLI-OTS, an off-the-shelf text language identification tool using the HeLI language identification method. The HeLI-OTS language identifier is equipped with language models for 200 languages and li…

Language Identification

bgGLUE: A Bulgarian General Language Understanding Evaluation Benchmark

2023-06-04 · Momchil Hardalov, Pepa Atanasova, Todor Mihaylov, Galia Angelova 외

We present bgGLUE(Bulgarian General Language Understanding Evaluation), a benchmark for evaluating language models on Natural Language Understanding (NLU) tasks in Bulgarian. Our benchmark includes NLU tasks targeting a …

Fact Checkingnamed-entity-recognitionNamed Entity RecognitionNatural Language Inference+3

A Publicly Available Cross-Platform Lemmatizer for Bulgarian

2015-06-13 · Grigor Iliev, Nadezhda Borisova, Elena Karashtranova, Dafina Kostadinova

Our dictionary-based lemmatizer for the Bulgarian language presented here is distributed as free software, publicly available to download and use under the GPL v3 license. The presented software is written entirely in Ja…

LemmatizationMORPH

Natural Language Processing Pipeline to Annotate Bulgarian Legislative Documents

2020-05-01 · LREC 2020 5 · Svetla Koeva, Nikola Obreshkov, Martin Yalamov

The paper presents the Bulgarian MARCELL corpus, part of a recently developed multilingual corpus representing the national legislation in seven European countries and the NLP pipeline that turns the web crawled data int…

Sentence

Bulgarian-English and English-Bulgarian Machine Translation: System Design and Evaluation

2017-09-01 · RANLP 2017 9 · Petya Osenova, Kiril Simov

The paper presents a deep factored machine translation (MT) system between English and Bulgarian languages in both directions. The MT system is hybrid. It consists of three main steps: (1) the source-language text is lin…

Machine TranslationTranslation