paper-with-me

홈 › Papers

Kawarith: an Arabic Twitter Corpus for Crisis Events

2021-04-01 · EACL (WANLP) 2021 4 · Alaa Alharbi, Mark Lee

Social media (SM) platforms such as Twitter provide large quantities of real-time data that can be leveraged during mass emergencies. Developing tools to support crisis-affected communities requires available datasets, which often do not exist for low resource languages. This paper introduces Kawarith a multi-dialect Arabic Twitter corpus for crisis events, comprising more than a million Arabic tweets collected during 22 crises that occurred between 2018 and 2020 and involved several types of hazard. Exploration of this content revealed the most discussed topics and information types, and the paper presents a labelled dataset from seven emergency events that serves as a gold standard for several tasks in crisis informatics research. Using annotated data from the same event, a BERT model is fine-tuned to classify tweets into different categories in the multi- label setting. Results show that BERT-based models yield good performance on this task even with small amounts of task-specific training data.

📄 PDF Abstract BibTeX

Code (2)

alaa-a-a/kawarith 공식 구현
alaa-a-a/multi-dialect-arabic-stop-words 공식 구현

Similar Papers 제목 키워드 기반

Building a Crisis Management Term Resource for Social Media: The Case of Floods and Protests

2014-05-01 · LREC 2014 5 · Irina Temnikova, Andrea Varga, Dogan Biyikli

Extracting information from social media is being currently exploited for a variety of tasks, including the recognition of emergency events in Twitter. This is done in order to supply Crisis Management agencies with addi…

DescriptiveInformation RetrievalManagementNamed Entity Recognition (NER)+1

DAICT: A Dialectal Arabic Irony Corpus Extracted from Twitter

2020-05-01 · LREC 2020 5 · Ines Abbes, Wajdi Zaghouani, Omaima El-Hardlo, Faten Ashour

Identifying irony in user-generated social media content has a wide range of applications; however to date Arabic content has received limited attention. To bridge this gap, this study builds a new open domain Arabic cor…

Arap-Tweet: A Large Multi-Dialect Twitter Corpus for Gender, Age and Language Variety Identification

2018-08-23 · LREC 2018 5 · Wajdi Zaghouani, Anis Charfi

In this paper, we present Arap-Tweet, which is a large-scale and multi-dialectal corpus of Tweets from 11 regions and 16 countries in the Arab world representing the major Arabic dialectal varieties. To build this corpus…

Author Profiling

An Arabic Twitter Corpus for Subjectivity and Sentiment Analysis

2014-05-01 · LREC 2014 5 · Eshrag Refaee, Verena Rieser

We present a newly collected data set of 8,868 gold-standard annotated Arabic feeds. The corpus is manually labelled for subjectivity and sentiment analysis (SSA) ( = 0:816). In addition, the corpus is annotated with a v…

Sentiment Analysis

Are We Ready for this Disaster? Towards Location Mention Recognition from Crisis Tweets

2020-12-01 · COLING 2020 8 · Reem Suwaileh, Muhammad Imran, Tamer Elsayed, Hassan Sajjad

The widespread usage of Twitter during emergencies has provided a new opportunity and timely resource to crisis responders for various disaster management tasks. Geolocation information of pertinent tweets is crucial for…

ArticlesManagement