An Arabic Twitter Corpus for Subjectivity and Sentiment Analysis
We present a newly collected data set of 8,868 gold-standard annotated Arabic feeds. The corpus is manually labelled for subjectivity and sentiment analysis (SSA) ( = 0:816). In addition, the corpus is annotated with a variety of motivated feature-sets that have previously shown positive impact on performance. The paper highlights issues posed by twitter as a genre, such as mixture of language varieties and topic-shifts. Our next step is to extend the current corpus, using online semi-supervised learning. A first sub-corpus will be released via the ELRA repository as part of this submission.
Code (0)
등록된 구현이 없습니다.
Tasks
Sentiment AnalysisSimilar Papers 제목 키워드 기반
Evaluating Distant Supervision for Subjectivity and Sentiment Analysis on Arabic Twitter Feeds
AWATIF: A Multi-Genre Corpus for Modern Standard Arabic Subjectivity and Sentiment Analysis
We present AWATIF, a multi-genre corpus of Modern Standard Arabic (MSA) labeled for subjectivity and sentiment analysis (SSA) at the sentence level. The corpus is labeled using both regular as well as crowd sourcing meth…
Opinion MiningSentenceSentiment AnalysisAn Arabic Tweets Sentiment Analysis Dataset (ATSAD) using Distant Supervision and Self Training
As the number of social media users increases, they express their thoughts, needs, socialise and publish their opinions reviews. For good social media sentiment analysis, good quality resources are needed, and the lack o…
8kArabic Sentiment AnalysisSentiment AnalysisSentiment Analysis of Arabic Tweets: Feature Engineering and A Hybrid Approach
Sentiment Analysis in Arabic is a challenging task due to the rich morphology of the language. Moreover, the task is further complicated when applied to Twitter data that is known to be highly informal and noisy. In this…
Feature EngineeringGeneral ClassificationSentiment Analysis