paper-with-me

홈 › Papers

What Is AI Safety? What Do We Want It to Be?

2025-05-05 · Jacqueline Harding, Cameron Domenico Kirk-Giannini

The field of AI safety seeks to prevent or reduce the harms caused by AI systems. A simple and appealing account of what is distinctive of AI safety as a field holds that this feature is constitutive: a research project falls within the purview of AI safety just in case it aims to prevent or reduce the harms caused by AI systems. Call this appealingly simple account The Safety Conception of AI safety. Despite its simplicity and appeal, we argue that The Safety Conception is in tension with at least two trends in the ways AI safety researchers and organizations think and talk about AI safety: first, a tendency to characterize the goal of AI safety research in terms of catastrophic risks from future systems; second, the increasingly popular idea that AI safety can be thought of as a branch of safety engineering. Adopting the methodology of conceptual engineering, we argue that these trends are unfortunate: when we consider what concept of AI safety it would be best to have, there are compelling reasons to think that The Safety Conception is the answer. Descriptively, The Safety Conception allows us to see how work on topics that have historically been treated as central to the field of AI safety is continuous with work on topics that have historically been treated as more marginal, like bias, misinformation, and privacy. Normatively, taking The Safety Conception seriously means approaching all efforts to prevent or mitigate harms from AI systems based on their merits rather than drawing arbitrary distinctions between them.

📄 PDF Abstract BibTeX arXiv:2505.02313

Code (0)

등록된 구현이 없습니다.

Tasks

Misinformation

Similar Papers 제목 키워드 기반

What Do AI-Generated Images Want?

2025-10-23 · Amanda Wasielewski arxiv

W.J.T. Mitchell's influential essay 'What do pictures want?' shifts the theoretical focus away from the interpretative act of understanding pictures and from the motivations of the humans who create them to the possibili…

Image Generation

Vision-Based Traffic Accident Detection and Anticipation: A Survey

2023-08-30 · Jianwu Fang, iahuan Qiao, Jianru Xue, Zhengguo Li

Traffic accident detection and anticipation is an obstinate road safety problem and painstaking efforts have been devoted. With the rapid growth of video data, Vision-based Traffic Accident Detection and Anticipation (na…

SurveyTraffic Accident Detection

A Case for AI Safety via Law

2023-07-31 · Jeffrey W. Johnston

How to make artificial intelligence (AI) systems safe and aligned with human values is an open research question. Proposed solutions tend toward relying on human intervention in uncertain situations, learning human value…

Designing for Human-Agent Alignment: Understanding what humans want from their agents

2024-04-04 · Nitesh Goyal, Minsuk Chang, Michael Terry

Our ability to build autonomous agents that leverage Generative AI continues to increase by the day. As builders and users of such agents it is unclear what parameters we need to align on before the agents start performi…

Ethics

Inverse Reward Design

2017-11-08 · NeurIPS 2017 12 · Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart Russell 외

Autonomous agents optimize the reward function we give them. What they don't know is how hard it is for us to design a reward function that actually captures what we want. When designing the reward, we might think of som…