paper-with-me

Papers

KeyXtract Twitter Model - An Essential Keywords Extraction Model for Twitter Designed using NLP Tools

2017-08-09 · Tharindu Weerasooriya, Nandula Perera, S. R. Liyanage

Since a tweet is limited to 140 characters, it is ambiguous and difficult for traditional Natural Language Processing (NLP) tools to analyse. This research presents KeyXtract which enhances the machine learning based Stanford CoreNLP Part-of-Speech (POS) tagger with the Twitter model to extract essential keywords from a tweet. The system was developed using rule-based parsers and two corpora. The data for the research was obtained from a Twitter profile of a telecommunication company. The system development consisted of two stages. At the initial stage, a domain specific corpus was compiled after analysing the tweets. The POS tagger extracted the Noun Phrases and Verb Phrases while the parsers removed noise and extracted any other keywords missed by the POS tagger. The system was evaluated using the Turing Test. After it was tested and compared against Stanford CoreNLP, the second stage of the system was developed addressing the shortcomings of the first stage. It was enhanced using Named Entity Recognition and Lemmatization. The second stage was also tested using the Turing test and its pass rate increased from 50.00% to 83.33%. The performance of the final system output was measured using the F1 score. Stanford CoreNLP with the Twitter model had an average F1 of 0.69 while the improved system had a F1 of 0.77. The accuracy of the system could be improved by using a complete domain specific corpus. Since the system used linguistic features of a sentence, it could be applied to other NLP tools.

📄 PDF Abstract BibTeX arXiv:1708.02912

Code (0)

등록된 구현이 없습니다.

Tasks

Lemmatizationmodelnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)POSSentence

Similar Papers 제목 키워드 기반

Thematic context vector association based on event uncertainty for Twitter

2023-04-04 · Vaibhav Khatavkar, Swapnil Mane, Parag Kulkarni

Keyword extraction is a crucial process in text mining. The extraction of keywords with respective contextual events in Twitter data is a big challenge. The challenging issues are mainly because of the informality in the…

Keyword ExtractionSarcasm Detection

Active Keyword Selection to Track Evolving Topics on Twitter

2022-09-22 · Sacha Lévy, Farimah Poursafaei, Kellin Pelrine, Reihaneh Rabbany

How can we study social interactions on evolving topics at a mass scale? Over the past decade, researchers from diverse fields such as economics, political science, and public health have often done this by querying Twit…

Active Learning

Natural Disaster Analysis using Satellite Imagery and Social-Media Data for Emergency Response Situations

2023-11-16 · Sukeerthi Mandyam, Shanmuga Priya MG, Shalini Suresh, Kavitha Srinivasan

Disaster Management is one of the most promising research areas because of its significant economic, environmental and social repercussions. This research focuses on analyzing different types of data (pre and post satell…

Management

Smart Crawling: A New Approach toward Focus Crawling from Twitter

2021-10-08 · Ahmad Khazaie, Nacéra Bennacer Seghouani, Francesca Bugiotti

Twitter is a social network that offers a rich and interesting source of information challenging to retrieve and analyze. Twitter data can be accessed using a REST API. The available operations allow retrieving tweets on…

Streaming Language-Specific Twitter Data with Optimal Keywords

2020-05-01 · LREC 2020 5 · Tim Kreutz, Walter Daelemans

The Twitter Streaming API has been used to create language-specific corpora with varying degrees of success. Selecting a filter of frequent yet distinct keywords for German resulted in a near-complete collection of Germa…