An Annotated Corpus for Sexism Detection in French Tweets
Social media networks have become a space where users are free to relate their opinions and sentiments which may lead to a large spreading of hatred or abusive messages which have to be moderated. This paper presents the first French corpus annotated for sexism detection composed of about 12,000 tweets. In a context of offensive content mediation on social media now regulated by European laws, we think that it is important to be able to detect automatically not only sexist content but also to identify if a message with a sexist content is really sexist (i.e. addressed to a woman or describing a woman or women in general) or is a story of sexism experienced by a woman. This point is the novelty of our annotation scheme. We also propose some preliminary results for sexism detection obtained with a deep learning approach. Our experiments show encouraging results.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
CoRoSeOf - An Annotated Corpus of Romanian Sexist and Offensive Tweets
This paper introduces CoRoSeOf, a large corpus of Romanian social media manually annotated for sexist and offensive language. We describe the annotation process of the corpus, provide initial analyses, and baseline class…
Binary ClassificationClassificationA French Corpus for Event Detection on Twitter
We present Event2018, a corpus annotated for event detection tasks, consisting of 38 million tweets in French (retweets excluded) including more than 130,000 tweets manually annotated by three annotators as related or un…
ArticlesEvent DetectionFrench Tweet Corpus for Automatic Stance Detection
The automatic stance detection task consists in determining the attitude expressed in a text toward a target (text, claim, or entity). This is a typical intermediate task for the fake news detection or analysis, which is…
Fake News DetectionStance DetectionCounter-TWIT: An Italian Corpus for Online Counterspeech in Ecological Contexts
This work describes the process of creating a corpus of Twitter conversations annotated for the presence of counterspeech in response to toxic speech related to axes of discrimination linked to sexism, racism and homopho…
Counterspeech DetectionTREMoLo-Tweets: A Multi-Label Corpus of French Tweets for Language Register Characterization
The casual, neutral, and formal language registers are highly perceptible in discourse productions. However, they are still poorly studied in Natural Language Processing (NLP), especially outside English, and for new tex…