paper-with-me

홈 › Papers

ROFF - A Romanian Twitter Dataset for Offensive Language

2021-09-01 · RANLP 2021 9 · Mihai Manolescu, Çağrı Çöltekin

This paper describes the annotation process of an offensive language data set for Romanian on social media. To facilitate comparable multi-lingual research on offensive language, the annotation guidelines follow some of the recent annotation efforts for other languages. The final corpus contains 5000 micro-blogging posts annotated by a large number of volunteer annotators. The inter-annotator agreement and the initial automatic discrimination results we present are in line with earlier annotation efforts.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RED: A Novel Dataset for Romanian Emotion Detection from Tweets

2021-09-01 · RANLP 2021 9 · Alexandra Ciobotaru, Liviu P. Dinu

In Romanian language there are some resources for automatic text comprehension, but for Emotion Detection, not lexicon-based, there are none. To cover this gap, we extracted data from Twitter and created the first datase…

BIG-bench Machine LearningClassificationOpinion MiningReading Comprehension+2

Arabic Offensive Language on Twitter: Analysis and Experiments

2020-04-05 · EACL (WANLP) 2021 4 · Hamdy Mubarak, Ammar Rashed, Kareem Darwish, Younes Samih 외

Detecting offensive language on Twitter has many applications ranging from detecting/predicting bullying to measuring polarization. In this paper, we focus on building a large Arabic offensive tweet dataset. We introduce…

CoRoSeOf - An Annotated Corpus of Romanian Sexist and Offensive Tweets

2022-06-01 · LREC 2022 6 · Diana Constantina Hoefels, Çağrı Çöltekin, Irina Diana Mădroane

This paper introduces CoRoSeOf, a large corpus of Romanian social media manually annotated for sexist and offensive language. We describe the annotation process of the corpus, provide initial analyses, and baseline class…

Binary ClassificationClassification

Emojis as Anchors to Detect Arabic Offensive Language and Hate Speech

2022-01-18 · Hamdy Mubarak, Sabit Hassan, Shammur Absar Chowdhury

We introduce a generic, language-independent method to collect a large percentage of offensive and hate tweets regardless of their topics or genres. We harness the extralinguistic information embedded in the emojis to co…

Cultural Vocal Bursts Intensity Prediction

Detecting Hate Speech and Offensive Language on Twitter using Machine Learning: An N-gram and TFIDF based Approach

2018-09-23 · Aditya Gaydhani, Vikrant Doma, Shrikant Kendre, Laxmi Bhagwat

Toxic online content has become a major issue in today's world due to an exponential increase in the use of internet by people of different cultures and educational background. Differentiating hate speech and offensive l…

Hate Speech Detection