paper-with-me

Papers

RECAST: Enabling User Recourse and Interpretability of Toxicity Detection Models with Interactive Visualization

2021-02-08 · Austin P Wright, Omar Shaikh, Haekyu Park, Will Epperson, Muhammed Ahmed, Stephane Pinel, Duen Horng Chau, Diyi Yang

With the widespread use of toxic language online, platforms are increasingly using automated systems that leverage advances in natural language processing to automatically flag and remove toxic comments. However, most automated systems -- when detecting and moderating toxic language -- do not provide feedback to their users, let alone provide an avenue of recourse for these users to make actionable changes. We present our work, RECAST, an interactive, open-sourced web tool for visualizing these models' toxic predictions, while providing alternative suggestions for flagged toxic language. Our work also provides users with a new path of recourse when using these automated moderation tools. RECAST highlights text responsible for classifying toxicity, and allows users to interactively substitute potentially toxic phrases with neutral alternatives. We examined the effect of RECAST via two large-scale user evaluations, and found that RECAST was highly effective at helping users reduce toxicity as detected through the model. Users also gained a stronger understanding of the underlying toxicity criterion used by black-box models, enabling transparency and recourse. In addition, we found that when users focus on optimizing language for these models instead of their own judgement (which is the implied incentive and goal of deploying automated models), these models cease to be effective classifiers of toxicity compared to human annotations. This opens a discussion for how toxicity detection models work and should work, and their effect on the future of online discourse.

📄 PDF Abstract BibTeX arXiv:2102.04427

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Recourse for reclamation: Chatting with generative language models

2024-03-21 · Jennifer Chien, Kevin R. McKee, Jackie Kay, William Isaac

Researchers and developers increasingly rely on toxicity scoring to moderate generative language model outputs, in settings such as customer service, information retrieval, and content generation. However, toxicity scori…

Information RetrievalLanguage ModelingLanguage ModellingRetrieval

Beyond Individualized Recourse: Interpretable and Interactive Summaries of Actionable Recourses

2020-09-15 · NeurIPS 2020 12 · Kaivalya Rawal, Himabindu Lakkaraju

As predictive models are increasingly being deployed in high-stakes decision-making, there has been a lot of interest in developing algorithms which can provide recourses to affected individuals. While developing such to…

counterfactualDecision Making

CARE: Coherent Actionable Recourse based on Sound Counterfactual Explanations

2021-08-18 · Peyman Rasouli, Ingrid Chieh Yu

Counterfactual explanation methods interpret the outputs of a machine learning model in the form of "what-if scenarios" without compromising the fidelity-interpretability trade-off. They explain how to obtain a desired p…

counterfactualCounterfactual ExplanationExplainable artificial intelligencetabular-classification

Target-confidence Recourse Using tSeTlin machines: TRUST

2026-06-17 · K. Darshana Abeyrathna, Sara El Mekkaoui, Nils Enric Canut Taugbøl, Anuja Vats arxiv

Counterfactual explanations are widely used to provide algorithmic recourse in high-stakes decision-making systems. Most existing methods seek the smallest change to an input that flips a model's decision. However, decis…

Low-Cost Algorithmic Recourse for Users With Uncertain Cost Functions

2021-11-01 · Prateek Yadav, Peter Hase, Mohit Bansal

People affected by machine learning model decisions may benefit greatly from access to recourses, i.e. suggestions about what features they could change to receive a more favorable decision from the model. Current approa…

Fairness