paper-with-me

홈 › Papers

AI red-teaming is a sociotechnical challenge: on values, labor, and harms

2024-12-12 · Tarleton Gillespie, Ryland Shaw, Mary L. Gray, Jina Suh

As generative AI technologies find more and more real-world applications, the importance of testing their performance and safety seems paramount. "Red-teaming" has quickly become the primary approach to test AI models--prioritized by AI companies, and enshrined in AI policy and regulation. Members of red teams act as adversaries, probing AI systems to test their safety mechanisms and uncover vulnerabilities. Yet we know far too little about this work or its implications. This essay calls for collaboration between computer scientists and social scientists to study the sociotechnical systems surrounding AI technologies, including the work of red-teaming, to avoid repeating the mistakes of the recent past. We highlight the importance of understanding the values and assumptions behind red-teaming, the labor arrangements involved, and the psychological impacts on red-teamers, drawing insights from the lessons learned around the work of content moderation.

📄 PDF Abstract BibTeX arXiv:2412.09751

Code (0)

등록된 구현이 없습니다.

Tasks

Red Teaming

Similar Papers 제목 키워드 기반

STAR: SocioTechnical Approach to Red Teaming Language Models

2024-06-17 · Laura Weidinger, John Mellor, Bernat Guillen Pegueroles, Nahema Marchal 외

This research introduces STAR, a sociotechnical framework that improves on current best practices for red teaming safety of large language models. STAR makes two key contributions: it enhances steerability by generating …

Red Teaming

Ask What Your Country Can Do For You: Towards a Public Red Teaming Model

2025-10-22 · Wm. Matthew Kennedy, Cigdem Patlak, Jayraj Dave, Blake Chambers 외 arxiv

AI systems have the potential to produce both benefits and harms, but without rigorous and ongoing adversarial evaluation, AI actors will struggle to assess the breadth and magnitude of the AI risk surface. Researchers f…

Red Teaming

Seeing Seeds Beyond Weeds: Green Teaming Generative AI for Beneficial Uses

2023-05-30 · Logan Stapleton, Jordan Taylor, Sarah Fox, Tongshuang Wu 외

Large generative AI models (GMs) like GPT and DALL-E are trained to generate content for general, wide-ranging purposes. GM content filters are generalized to filter out content which has a risk of harm in many cases, e.…

Red Teaming

A Framework for Automated Measurement of Responsible AI Harms in Generative AI Applications

2023-10-26 · Ahmed Magooda, Alec Helyar, Kyle Jackson, David Sullivan 외

We present a framework for the automated measurement of responsible AI (RAI) metrics for large language models (LLMs) and associated products and services. Our framework for automatically measuring harms from LLMs builds…

Ontologies in Design: How Imagining a Tree Reveals Possibilites and Assumptions in Large Language Models

2025-04-03 · Nava Haghighi, Sunny Yu, James Landay, Daniela Rosner

Amid the recent uptake of Generative AI, sociotechnical scholars and critics have traced a multitude of resulting harms, with analyses largely focused on values and axiology (e.g., bias). While value-based analyses are c…