paper-with-me

홈 › Papers

OffendES: A New Corpus in Spanish for Offensive Language Research

2021-09-01 · RANLP 2021 9 · Flor Miriam Plaza-del-Arco, Arturo Montejo-Ráez, L. Alfonso Ureña-López, María-Teresa Martín-Valdivia

Offensive language detection and analysis has become a major area of research in Natural Language Processing. The freedom of participation in social media has exposed online users to posts designed to denigrate, insult or hurt them according to gender, race, religion, ideology, or other personal characteristics. Focusing on young influencers from the well-known social platforms of Twitter, Instagram, and YouTube, we have collected a corpus composed of 47,128 Spanish comments manually labeled on offensive pre-defined categories. A subset of the corpus attaches a degree of confidence to each label, so both multi-class classification and multi-output regression studies are possible. In this paper, we introduce the corpus, discuss its building process, novelties, and some preliminary experiments with it to serve as a baseline for the research community.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-class Classificationregression

Similar Papers 제목 키워드 기반

SHARE: A Lexicon of Harmful Expressions by Spanish Speakers

2022-06-01 · LREC 2022 6 · Flor Miriam Plaza-del-Arco, Ana Belén Parras Portillo, Pilar López Úbeda, Beatriz Gil 외

In this paper we present SHARE, a new lexical resource with 10,125 offensive terms and expressions collected from Spanish speakers. We retrieve this vocabulary using an existing chatbot developed to engage a conversation…

Chatbot

Automatic Detection of Offensive Language in Social Media: Defining Linguistic Criteria to build a Mexican Spanish Dataset

2020-05-01 · LREC 2020 5 · Mar{\'\i}a Jos{\'e} D{\'\i}az-Torres, Paulina Alej Mor{\'a}n-M{\'e}ndez, ra, Luis Villasenor-Pineda 외

Phenomena such as bullying, homophobia, sexism and racism have transcended to social networks, motivating the development of tools for their automatic detection. The challenge becomes greater for languages rich in popula…

Abusive Language

Predicting the Type and Target of Offensive Social Media Posts in Marathi

2022-11-22 · Marcos Zampieri, Tharindu Ranasinghe, Mrinal Chaudhari, Saurabh Gaikwad 외

The presence of offensive language on social media is very common motivating platforms to invest in strategies to make communities safer. This includes developing robust machine learning systems capable of recognizing of…

Language Identification

HateBR: A Large Expert Annotated Corpus of Brazilian Instagram Comments for Offensive Language and Hate Speech Detection

2021-03-27 · LREC 2022 6 · Francielle Alves Vargas, Isabelle Carvalho, Fabiana Rodrigues de Góes, Fabrício Benevenuto 외

Due to the severity of the social media offensive and hateful comments in Brazil, and the lack of research in Portuguese, this paper provides the first large-scale expert annotated corpus of Brazilian Instagram comments …

BIG-bench Machine LearningBinary ClassificationHate Speech Detection

EmoEvent: A Multilingual Emotion Corpus based on different Events

2020-05-01 · LREC 2020 5 · Flor Miriam Plaza del Arco, Carlo Strapparava, L. Alfonso Urena Lopez, Maite Martin

In recent years emotion detection in text has become more popular due to its potential applications in fields such as psychology, marketing, political science, and artificial intelligence, among others. While opinion min…

MarketingOpinion Mining