paper-with-me

홈 › Papers

ToxiREX: A Dataset on Toxic REasoning in ConteXt

2026-06-26 · Stefan F. Schouten, Ilia Markov, Piek Vossen arxiv

We introduce a new, contextual, multilingual dataset called ToxiREX: Toxic REasoning in ConteXt. The dataset consists of threads of Reddit comments and structured characterizations of what the comments imply, following a systematic toxic reasoning schema developed in a previous paper. Using the schema allows us to capture and explain implicit and context-dependent toxicity, while supporting mappings to existing toxicity taxonomies. The dataset includes comments in six languages (English, Arabic, Turkish, Spanish, German, and Dutch), collected from posts connected to specific major events (e.g. the 2023 Turkey earthquakes; the Russian invasion of Ukraine). We describe the context-preserving preprocessing of the threads. We create a training set of 125 thousand comments which is annotated by a commercially available LLM, and a test set of just under three thousand comments that is annotated by native speakers. We show that apparent disagreements in the test set annotations often reflect defensible alternative interpretations rather than noise. Finally, we provide baseline results by prompting and fine-tuning language models. To produce these results, we develop evaluation strategies for our hierarchical, schema-based predictions. While models perform better than random, there remains a lot of room for improvement, showing the task to be challenging. ToxiREX is the first dataset to simultaneously incorporate multiple languages, conversational context, and implicit toxicity, while using the toxic reasoning schema for rich, structured annotations. Dataset available at: https://github.com/cltl/toxirex

📄 PDF Abstract BibTeX arXiv:2606.27981

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CoTox: Chain-of-Thought-Based Molecular Toxicity Reasoning and Prediction

2025-08-05 · Jueon Park, Yein Park, Minju Song, Soyon Park 외 arxiv

Drug toxicity remains a major challenge in pharmaceutical development. Recent machine learning models have improved in silico toxicity prediction, but their reliance on annotated data and lack of interpretability limit t…

Just KIDDIN: Knowledge Infusion and Distillation for Detection of INdecent Memes

2024-11-19 · Rahul Garg, Trilok Padhi, Hemang Jain, Ugur Kursuncu 외

Toxicity identification in online multimodal environments remains a challenging task due to the complexity of contextual connections across modalities (e.g., textual and visual). In this paper, we propose a novel framewo…

Knowledge DistillationKnowledge Graphs

Context Sensitivity Estimation in Toxicity Detection

2021-08-01 · ACL (WOAH) 2021 8 · Alexandros Xenos, John Pavlopoulos, Ion Androutsopoulos

User posts whose perceived toxicity depends on the conversational context are rare in current toxicity detection datasets. Hence, toxicity detectors trained on current datasets will also disregard context, making the det…

Sensitivity

Toxicity Detection can be Sensitive to the Conversational Context

2021-11-19 · Alexandros Xenos, John Pavlopoulos, Ion Androutsopoulos, Lucas Dixon 외

User posts whose perceived toxicity depends on the conversational context are rare in current toxicity detection datasets. Hence, toxicity detectors trained on existing datasets will also tend to disregard context, makin…

Data AugmentationKnowledge Distillation

ToxReason: A Benchmark for Mechanistic Chemical Toxicity Reasoning via Adverse Outcome Pathway

2026-04-07 · Jueon Park, Wonjune Jang, Chanhwi Kim, Yein Park 외 arxiv

Recent advances in large language models (LLMs) have enabled molecular reasoning for property prediction. However, toxicity arises from complex biological mechanisms beyond chemical structure, necessitating mechanistic r…