paper-with-me

Papers

Risk Reporting for Developers' Internal AI Model Use

2026-04-27 · Oscar Delaney, Sambhav Maheshwari, Joe O'Brien, Theo Bearman, Oliver Guest arxiv

Frontier AI companies first deploy their most advanced models internally, for weeks or months of safety testing, evaluation, and iteration, before a possible public release. For example, Anthropic recently developed a new class of model with advanced cyberoffense-relevant capabilities, Mythos Preview, which was available internally for at least six weeks before it was publicly announced. This internal use creates risks that external deployment frameworks may fail to address. Legal frameworks, notably California's Transparency in Frontier Artificial Intelligence Act (SB 53), New York's Responsible AI Safety And Education (RAISE) Act, and the EU's General-Purpose AI Code of Practice, all discuss risks from internal AI use. They require frontier developers to make and implement plans for how to manage risks from internal use, and to produce internal use risk reports describing their safeguards and any residual risks. This guide provides a harmonized standard for companies to produce internal use risk reports suitable for all three regulatory frameworks. It is addressed primarily to evaluation and safety teams at frontier AI developers, and secondarily to regulators and auditors seeking to understand what good reporting looks like. Given the pace of AI R&D automation and the limited external visibility into how companies use their most capable models internally, regular and detailed risk reporting may be one of the few mechanisms available to ensure that the risks from internal AI use are identified and managed before they materialize. Whenever a substantially more capable or riskier model is deployed internally, the developer should create a risk report and argue why the model is safe to deploy. We structure the reporting framework around two threat vectors -- autonomous AI misbehavior and insider threats -- and three risk factors for each: means, motive, and opportunity.

📄 PDF Abstract BibTeX arXiv:2604.24966

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The AI Model Risk Catalog: What Developers and Researchers Miss About Real-World AI Harms

2025-08-21 · Pooja S. B. Rao, Sanja Šćepanović, Dinesh Babu Jayagopi, Mauro Cherubini 외 arxiv

We analyzed nearly 460,000 AI model cards from Hugging Face to examine how developers report risks. From these, we extracted around 3,000 unique risk mentions and built the \emph{AI Model Risk Catalog}. We compared these…

Responsible Reporting for Frontier AI Development

2024-04-03 · Noam Kolt, Markus Anderljung, Joslyn Barnhart, Asher Brass 외

Mitigating the risks from frontier AI systems requires up-to-date and reliable information about those systems. Organizations that develop and deploy frontier systems have significant access to such information. By repor…

Management

STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports

2025-08-13 · Tegan McCaslin, Jide Alaga, Samira Nedungadi, Seth Donoughe 외 arxiv

Evaluations of dangerous AI capabilities are important for managing catastrophic risks. Public transparency into these evaluations - including what they test, how they are conducted, and how their results inform decision…

FLARE-AI: Flaw Reporting for AI

2026-06-30 · Shayne Longpre, Elaine Zhu, Carson Ezell, Avijit Ghosh 외 arxiv

Flaw reporting for deployed AI systems is fundamental to identifying system failures and improving AI safety. Yet the AI reporting ecosystem is fragmented: researchers who identify flaws often do not know what or where t…

Defending Compute Thresholds Against Legal Loopholes

2025-01-03 · Matteo Pistillo, Pablo Villalobos

Existing legal frameworks on AI rely on training compute thresholds as a proxy to identify potentially-dangerous AI models and trigger increased regulatory attention. In the United States, Section 4.2(a) of Executive Ord…