paper-with-me

Papers

Automatically Annotated Turkish Corpus for Named Entity Recognition and Text Categorization using Large-Scale Gazetteers

2017-02-08 · H. Bahadir Sahin, Caglar Tirkaz, Eray Yildiz, Mustafa Tolga Eren, Ozan Sonmez

Turkish Wikipedia Named-Entity Recognition and Text Categorization (TWNERTC) dataset is a collection of automatically categorized and annotated sentences obtained from Wikipedia. We constructed large-scale gazetteers by using a graph crawler algorithm to extract relevant entity and domain information from a semantic knowledge base, Freebase. The constructed gazetteers contains approximately 300K entities with thousands of fine-grained entity types under 77 different domains. Since automated processes are prone to ambiguity, we also introduce two new content specific noise reduction methodologies. Moreover, we map fine-grained entity types to the equivalent four coarse-grained types: person, loc, org, misc. Eventually, we construct six different dataset versions and evaluate the quality of annotations by comparing ground truths from human annotators. We make these datasets publicly available to support studies on Turkish named-entity recognition (NER) and text categorization (TC).

📄 PDF Abstract BibTeX arXiv:1702.02363

Code (0)

등록된 구현이 없습니다.

Tasks

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NERText Categorization

Similar Papers 제목 키워드 기반

Named Entity Recognition on Turkish Tweets

2014-05-01 · LREC 2014 5 · Dilek K{\"u}{\c{c}}{\"u}k, Guillaume Jacquet, Ralf Steinberger

Various recent studies show that the performance of named entity recognition (NER) systems developed for well-formed text types drops significantly when applied to tweets. The only existing study for the highly inflected…

Articlesnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

A Twitter Corpus for Named Entity Recognition in Turkish

2022-06-01 · LREC 2022 6 · Buse Çarık, Reyyan Yeniterzi

This paper introduces a new Turkish Twitter Named Entity Recognition dataset. The dataset, which consists of 5000 tweets from a year-long period, was labeled by multiple annotators with a high agreement score. The datase…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

Transfer-based Enrichment of a Hungarian Named Entity Dataset

2021-09-01 · RANLP 2021 9 · Attila Novák, Borbála Novák

In this paper, we present a major update to the first Hungarian named entity dataset, the Szeged NER corpus. We used zero-shot cross-lingual transfer to initialize the enrichment of entity types annotated in the corpus u…

Cross-Lingual TransferNERZero-Shot Cross-Lingual Transfer

Experiments to Improve Named Entity Recognition on Turkish Tweets

2014-10-31 · WS 2014 4 · Dilek Küçük, Ralf Steinberger

Social media texts are significant information sources for several application areas including trend analysis, event monitoring, and opinion mining. Unfortunately, existing solutions for tasks such as named entity recogn…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Opinion Mining

UNER: Universal Named-Entity RecognitionFramework

2020-10-23 · Diego Alves, Tin Kuculo, Gabriel Amaral, Gaurish Thakkar 외

We introduce the Universal Named-Entity Recognition (UNER)framework, a 4-level classification hierarchy, and the methodology that isbeing adopted to create the first multilingual UNER corpus: the SETimesparallel corpus a…

Knowledge Graphsnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1