paper-with-me

Papers

Arap-Tweet: A Large Multi-Dialect Twitter Corpus for Gender, Age and Language Variety Identification

2018-08-23 · LREC 2018 5 · Wajdi Zaghouani, Anis Charfi

In this paper, we present Arap-Tweet, which is a large-scale and multi-dialectal corpus of Tweets from 11 regions and 16 countries in the Arab world representing the major Arabic dialectal varieties. To build this corpus, we collected data from Twitter and we provided a team of experienced annotators with annotation guidelines that they used to annotate the corpus for age categories, gender, and dialectal variety. During the data collection effort, we based our search on distinctive keywords that are specific to the different Arabic dialects and we also validated the location using Twitter API. In this paper, we report on the corpus data collection and annotation efforts. We also present some issues that we encountered during these phases. Then, we present the results of the evaluation performed to ensure the consistency of the annotation. The provided corpus will enrich the limited set of available language resources for Arabic and will be an invaluable enabler for developing author profiling tools and NLP tools for Arabic.

📄 PDF Abstract BibTeX arXiv:1808.07674

Code (0)

등록된 구현이 없습니다.

Tasks

Author Profiling

Similar Papers 제목 키워드 기반

A Fine-Grained Annotated Multi-Dialectal Arabic Corpus

2019-09-01 · RANLP 2019 9 · Anis Charfi, Wajdi Zaghouani, Syed Hassan Mehdi, Esraa Mohamed

We present ARAP-Tweet 2.0, a corpus of 5 million dialectal Arabic tweets and 50 million words of about 3000 Twitter users from 17 Arab countries. Compared to the first version, the new corpus has significant improvements…

Arabic Offensive Language on Twitter: Analysis and Experiments

2020-04-05 · EACL (WANLP) 2021 4 · Hamdy Mubarak, Ammar Rashed, Kareem Darwish, Younes Samih 외

Detecting offensive language on Twitter has many applications ranging from detecting/predicting bullying to measuring polarization. In this paper, we focus on building a large Arabic offensive tweet dataset. We introduce…

An Algerian Corpus and an Annotation Platform for Opinion and Emotion Analysis

2020-05-01 · LREC 2020 5 · Leila Moudjari, Karima Akli-Astouati, Farah Benamara

In this paper, we address the lack of resources for opinion and emotion analysis related to North African dialects, targeting Algerian dialect. We present TWIFIL (TWItter proFILing) a collaborative annotation platform fo…

Author ProfilingEmotion Recognition

TKLBLIIR: Detecting Twitter Paraphrases with TweetingJay

2015-06-01 · SEMEVAL 2015 6 · Mladen Karan, Goran Glava{\v{s}}, Jan {\v{S}}najder, Bojana Dalbelo Ba{\v{s}}i{\'c} 외
Information RetrievalMachine TranslationSemantic Textual SimilarityWord Alignment

Acquiring Predicate Paraphrases from News Tweets

2017-08-01 · SEMEVAL 2017 8 · Vered Shwartz, Gabriel Stanovsky, Ido Dagan

We present a simple method for ever-growing extraction of predicate paraphrases from news headlines in Twitter. Analysis of the output of ten weeks of collection shows that the accuracy of paraphrases with different supp…

Natural Language InferenceQuestion Answering