Corpus Creation and Analysis for Named Entity Recognition in Telugu-English Code-Mixed Social Media Data
Named Entity Recognition(NER) is one of the important tasks in Natural Language Processing(NLP) and also is a subtask of Information Extraction. In this paper we present our work on NER in Telugu-English code-mixed social media data. Code-Mixing, a progeny of multilingualism is a way in which multilingual people express themselves on social media by using linguistics units from different languages within a sentence or speech context. Entity Extraction from social media data such as tweets(twitter) is in general difficult due to its informal nature, code-mixed data further complicates the problem due to its informal, unstructured and incomplete information. We present a Telugu-English code-mixed corpus with the corresponding named entity tags. The named entities used to tag data are Person({}Per{'}), Organization({}Org{'}) and Location({`}Loc{'}). We experimented with the machine learning models Conditional Random Fields(CRFs), Decision Trees and BiLSTMs on our corpus which resulted in a F1-score of 0.96, 0.94 and 0.95 respectively.
Code (0)
등록된 구현이 없습니다.
Tasks
Entity Extraction using GANnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NERSentenceTAGSimilar Papers 제목 키워드 기반
Domain Adaptation for Named Entity Recognition Using CRFs
In this paper we explain how we created a labelled corpus in English for a Named Entity Recognition (NER) task from multi-source and multi-domain data, for an industrial partner. We explain the specificities of this corp…
Domain Adaptationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1An Open Corpus for Named Entity Recognition in Historic Newspapers
The availability of openly available textual datasets ({``}corpora{''}) with highly accurate manual annotations ({``}gold standard{''}) of named entities (e.g. persons, locations, organizations, etc.) is crucial in the t…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Creation of Corpus and Analysis in Code-Mixed Kannada-English Social Media Data for POS Tagging
Part-of-Speech (POS) is one of the essential tasks for many Natural Language Processing (NLP) applications. There has been a significant amount of work done in POS tagging for resource-rich languages. POS tagging is an e…
coreference-resolutionCoreference Resolutionnamed-entity-recognitionNamed Entity Recognition+5Named Entity Recognition on Turkish Tweets
Various recent studies show that the performance of named entity recognition (NER) systems developed for well-formed text types drops significantly when applied to tweets. The only existing study for the highly inflected…
Articlesnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1