paper-with-me

홈 › Papers

ConvAbuse: Data, Analysis, and Benchmarks for Nuanced Abuse Detection in Conversational AI

2021-09-20 · Amanda Cercas Curry, Gavin Abercrombie, Verena Rieser

We present the first English corpus study on abusive language towards three conversational AI systems gathered "in the wild": an open-domain social bot, a rule-based chatbot, and a task-based system. To account for the complexity of the task, we take a more `nuanced' approach where our ConvAI dataset reflects fine-grained notions of abuse, as well as views from multiple expert annotators. We find that the distribution of abuse is vastly different compared to other commonly used datasets, with more sexually tinted aggression towards the virtual persona of these systems. Finally, we report results from bench-marking existing models against this data. Unsurprisingly, we find that there is substantial room for improvement with F1 scores below 90%.

📄 PDF Abstract BibTeX arXiv:2109.09483

Code (1)

amandacurry/convabuse 공식 구현

Tasks

Abuse DetectionAbusive LanguageChatbot

Similar Papers 제목 키워드 기반

ConvAbuse: Data, Analysis, and Benchmarks for Nuanced Detection in Conversational AI

2021-11-01 · EMNLP 2021 11 · Amanda Cercas Curry, Gavin Abercrombie, Verena Rieser

We present the first English corpus study on abusive language towards three conversational AI systems gathered ‘in the wild’: an open-domain social bot, a rule-based chatbot, and a task-based system. To account for the c…

Abusive LanguageChatbot

Ensemble Diversity Optimization for Subjective Supervision

2026-07-09 · Xia Cui, Ziyi Huang, N. R. Abeynayake arxiv

Subjective NLP tasks often exhibit systematic annotator disagreement, requiring models that represent uncertainty rather than collapse it. We introduce Ensemble Diversity Optimization (EDO), a prediction-space framework …

Large Language Models in the Abuse Detection Pipeline

2026-03-31 · Suraj Kath, Sanket Badhe, Preet Shah, Ashwin Sampathkumar 외 arxiv

Online abuse has grown increasingly complex, spanning toxic language, harassment, manipulation, and fraudulent behavior. Traditional machine-learning approaches dependent on static classifiers and labor-intensive labelin…

Adversarial RobustnessExplanation Generation

``Are you kidding me?'': Detecting Unpalatable Questions on Reddit

2021-04-01 · EACL 2021 2 · Sunyam Bagga, Andrew Piper, Derek Ruths

Abusive language in online discourse negatively affects a large number of social media users. Many computational methods have been proposed to address this issue of online abuse. The existing work, however, tends to focu…

Abusive Language

Addressing Antisocial Behavior in Multi-Party Dialogs Through Multimodal Representation Learning

2025-10-20 · Hajar Bakarou, Mohamed Sinane El Messoussi, Anaïs Ollagnier arxiv

Antisocial behavior (ASB) on social media -- including hate speech, harassment, and cyberbullying -- poses growing risks to platform safety and societal well-being. Prior research has focused largely on networks such as …

Representation Learning