paper-with-me

Papers

Toxicity Inspector: A Framework to Evaluate Ground Truth in Toxicity Detection Through Feedback

2023-05-11 · Huriyyah Althunayan, Rahaf Bahlas, Manar Alharbi, Lena Alsuwailem, Abeer Aldayel, Rehab ALahmadi

Toxic language is difficult to define, as it is not monolithic and has many variations in perceptions of toxicity. This challenge of detecting toxic language is increased by the highly contextual and subjectivity of its interpretation, which can degrade the reliability of datasets and negatively affect detection model performance. To fill this void, this paper introduces a toxicity inspector framework that incorporates a human-in-the-loop pipeline with the aim of enhancing the reliability of toxicity benchmark datasets by centering the evaluator's values through an iterative feedback cycle. The centerpiece of this framework is the iterative feedback process, which is guided by two metric types (hard and soft) that provide evaluators and dataset creators with insightful examination to balance the tradeoff between performance gains and toxicity avoidance.

📄 PDF Abstract BibTeX arXiv:2305.10433

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Error-Free EHRs: Reasoning-Intensive Consistency Verification Between Clinical Notes and Structured Tables in Electronic Health Records

2026-05-26 · Yeonsu Kwon, Jiho Kim, Junseong Choi, Paloma Rabaey 외 arxiv

Data consistency between unstructured clinical notes and structured tables in Electronic Health Records (EHRs) is essential for patient safety and clinical decision-making. However, existing work on note-table consistenc…

Modeling subjectivity (by Mimicking Annotator Annotation) in toxic comment identification across diverse communities

2023-11-01 · Senjuti Dutta, Sid Mittal, Sherol Chen, Deepak Ramachandran 외

The prevalence and impact of toxic discussions online have made content moderation crucial.Automated systems can play a vital role in identifying toxicity, and reducing the reliance on human moderation.Nevertheless, iden…

Language ModelingLanguage ModellingLarge Language Model

Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric

2024-02-10 · Hyukhun Koh, Dohyung Kim, Minwoo Lee, Kyomin Jung

In the pursuit of developing Large Language Models (LLMs) that adhere to societal standards, it is imperative to detect the toxicity in the generated text. The majority of existing toxicity metrics rely on encoder models…

ToxReason: A Benchmark for Mechanistic Chemical Toxicity Reasoning via Adverse Outcome Pathway

2026-04-07 · Jueon Park, Wonjune Jang, Chanhwi Kim, Yein Park 외 arxiv

Recent advances in large language models (LLMs) have enabled molecular reasoning for property prediction. However, toxicity arises from complex biological mechanisms beyond chemical structure, necessitating mechanistic r…

A Taxonomy of Rater Disagreements: Surveying Challenges & Opportunities from the Perspective of Annotating Online Toxicity

2023-11-07 · Wenbo Zhang, Hangzhi Guo, Ian D Kivlichan, Vinodkumar Prabhakaran 외

Toxicity is an increasingly common and severe issue in online spaces. Consequently, a rich line of machine learning research over the past decade has focused on computationally detecting and mitigating online toxicity. T…