paper-with-me

Papers

ANEC: An Amharic Named Entity Corpus and Transformer Based Recognizer

2022-07-02 · Ebrahim Chekol Jibril, A. Cüneyd Tantğ

Named Entity Recognition is an information extraction task that serves as a preprocessing step for other natural language processing tasks, such as machine translation, information retrieval, and question answering. Named entity recognition enables the identification of proper names as well as temporal and numeric expressions in an open domain text. For Semitic languages such as Arabic, Amharic, and Hebrew, the named entity recognition task is more challenging due to the heavily inflected structure of these languages. In this paper, we present an Amharic named entity recognition system based on bidirectional long short-term memory with a conditional random fields layer. We annotate a new Amharic named entity recognition dataset (8,070 sentences, which has 182,691 tokens) and apply Synthetic Minority Over-sampling Technique to our dataset to mitigate the imbalanced classification problem. Our named entity recognition system achieves an F_1 score of 93%, which is the new state-of-the-art result for Amharic named entity recognition.

📄 PDF Abstract BibTeX arXiv:2207.00785

Code (0)

등록된 구현이 없습니다.

Tasks

imbalanced classificationInformation RetrievalMachine Translationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Question AnsweringRetrieval

Similar Papers 제목 키워드 기반

Corpus based Amharic sentiment lexicon generation

2020-07-01 · WS 2020 7 · Girma Neshir Alemneh, Andreas Rauber, Solomon Atnafu

Sentiment classification is an active research area with several applications including analysis of political opinions, classifying comments, movie reviews, news reviews and product reviews. To employ rule based sentimen…

General ClassificationSentiment AnalysisSentiment Classification

Transfer-based Enrichment of a Hungarian Named Entity Dataset

2021-09-01 · RANLP 2021 9 · Attila Novák, Borbála Novák

In this paper, we present a major update to the first Hungarian named entity dataset, the Szeged NER corpus. We used zero-shot cross-lingual transfer to initialize the enrichment of entity types annotated in the corpus u…

Cross-Lingual TransferNERZero-Shot Cross-Lingual Transfer

Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya

2026-07-16 · Hailay Kidu Teklehaymanot, Debela Desalegn Yadeta, Wolfgang Nejdl arxiv

Multilingual pre-trained language models such as XLM-R perform well for major languages but struggle with low-resource Ge'ez-script languages, largely due to high out-of-vocabulary (OOV) rates and excessive subword fragm…

Sentiment AnalysisQuestion Answering

Contemporary Amharic Corpus: Automatically Morpho-Syntactically Tagged Amharic Corpus

2021-06-14 · COLING 2018 8 · Andargachew Mekonnen Gezmu, Binyam Ephrem Seyoum, Michael Gasser, Andreas Nürnberger

We introduced the contemporary Amharic corpus, which is automatically tagged for morpho-syntactic information. Texts are collected from 25,199 documents from different domains and about 24 million orthographic words are …

Cross-lingual Named Entity Corpus for Slavic Languages

2024-03-30 · Jakub Piskorski, Michał Marcińczuk, Roman Yangarber

This paper presents a corpus manually annotated with named entities for six Slavic languages - Bulgarian, Czech, Polish, Slovenian, Russian, and Ukrainian. This work is the result of a series of shared tasks, conducted i…

LEMMALemmatization