paper-with-me

Papers

Output Scouting: Auditing Large Language Models for Catastrophic Responses

2024-10-04 · Andrew Bell, Joao Fonseca

Recent high profile incidents in which the use of Large Language Models (LLMs) resulted in significant harm to individuals have brought about a growing interest in AI safety. One reason LLM safety issues occur is that models often have at least some non-zero probability of producing harmful outputs. In this work, we explore the following scenario: imagine an AI safety auditor is searching for catastrophic responses from an LLM (e.g. a "yes" responses to "can I fire an employee for being pregnant?"), and is able to query the model a limited number times (e.g. 1000 times). What is a strategy for querying the model that would efficiently find those failure responses? To this end, we propose output scouting: an approach that aims to generate semantically fluent outputs to a given prompt matching any target probability distribution. We then run experiments using two LLMs and find numerous examples of catastrophic responses. We conclude with a discussion that includes advice for practitioners who are looking to implement LLM auditing for catastrophic responses. We also release an open-source toolkit (https://github.com/joaopfonseca/outputscouting) that implements our auditing framework using the Hugging Face transformers library.

📄 PDF Abstract BibTeX arXiv:2410.05305

Code (1)

joaopfonseca/outputscouting 공식 구현

Similar Papers 제목 키워드 기반

Automatically Auditing Large Language Models via Discrete Optimization

2023-03-08 · Erik Jones, Anca Dragan, aditi raghunathan, Jacob Steinhardt

Auditing large language models for unexpected behaviors is critical to preempt catastrophic deployments, yet remains challenging. In this work, we cast auditing as an optimization problem, where we automatically search f…

CALM: Curiosity-Driven Auditing for Large Language Models

2025-01-06 · Xiang Zheng, Longxiang Wang, Yi Liu, Xingjun Ma 외

Auditing Large Language Models (LLMs) is a crucial and challenging task. In this study, we focus on auditing black-box LLMs without access to their parameters, only to the provided service. We treat this type of auditing…

NFL Career Success as Predicted by NFL Scouting Combine

2023-03-10 · Brian Szekely, Christian Sinnott, Savannah Halow, Gregory Ryan

The National Football League (NFL) Scouting Combine serves as a tool to evaluate the skills of prospective players and assess their readiness to play in the NFL. The development of machine learning brings new opportuniti…

Trouble with the Curve: Predicting Future MLB Players Using Scouting Reports

2019-10-21 · Jacob Danovitch

In baseball, a scouting report profiles a player's characteristics and traits, usually intended for use in player valuation. This work presents a first-of-its-kind dataset of almost 10,000 scouting reports for minor leag…

Articles

Hunt Globally: Wide Search AI Agents for Drug Asset Scouting in Investing, Business Development, and Competitive Intelligence

2026-02-16 · Vlad Vinogradov, Alisa Vinogradova, Luba Greenwood, Ilya Yasny 외 arxiv

Bio-pharmaceutical innovation has shifted: many new drug assets now originate outside the United States and are disclosed primarily via regional, non-English channels. Recent data suggests over 85% of patent filings orig…