paper-with-me

홈 › Papers

Foundational Challenges in Assuring Alignment and Safety of Large Language Models

2024-04-15 · Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, Benjamin L. Edelman, Zhaowei Zhang, Mario Günther, Anton Korinek, Jose Hernandez-Orallo, Lewis Hammond, Eric Bigelow, Alexander Pan, Lauro Langosco, Tomasz Korbak, Heidi Zhang, Ruiqi Zhong, Seán Ó hÉigeartaigh, Gabriel Recchia, Giulio Corsi, Alan Chan, Markus Anderljung, Lilian Edwards, Aleksandar Petrov, Christian Schroeder de Witt, Sumeet Ramesh Motwan, Yoshua Bengio, Danqi Chen, Philip H. S. Torr, Samuel Albanie, Tegan Maharaj, Jakob Foerster, Florian Tramer, He He, Atoosa Kasirzadeh, Yejin Choi, David Krueger

This work identifies 18 foundational challenges in assuring the alignment and safety of large language models (LLMs). These challenges are organized into three different categories: scientific understanding of LLMs, development and deployment methods, and sociotechnical challenges. Based on the identified challenges, we pose $200+$ concrete research questions.

📄 PDF Abstract BibTeX arXiv:2404.09932

Code (1)

thuccslab/figstep

Similar Papers 제목 키워드 기반

Assuring the Safety of Reinforcement Learning Components: AMLAS-RL

2025-07-08 · Calum Corrie Imrie, Ioannis Stefanakos, Sepeedeh Shahbeigi, Richard Hawkins 외 arxiv

The rapid advancement of machine learning (ML) has led to its increasing integration into cyber-physical systems (CPS) across diverse domains. While CPS offer powerful capabilities, incorporating ML components introduces…

Reinforcement Learning

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning

2025-03-05 · Borong Zhang, Yuhao Zhang, Jiaming Ji, Yingshan Lei 외

Vision-language-action models (VLAs) show potential as generalist robot policies. However, these models pose extreme safety challenges during real-world deployment, including the risk of harm to the environment, the robo…

Safe Reinforcement LearningSafety AlignmentVision-Language-Action

Testing and verification of neural-network-based safety-critical control software: A systematic literature review

2019-10-05 · Jin Zhang, Jingyue Li

Context: Neural Network (NN) algorithms have been successfully adopted in a number of Safety-Critical Cyber-Physical Systems (SCCPSs). Testing and Verification (T&V) of NN-based control software in safety-critical domain…

Systematic Literature Review

Landscape of AI safety concerns -- A methodology to support safety assurance for AI-based autonomous systems

2024-12-18 · Ronald Schnitzer, Lennart Kilian, Simon Roessner, Konstantinos Theodorou 외

Artificial Intelligence (AI) has emerged as a key technology, driving advancements across a range of applications. Its integration into modern autonomous systems requires assuring safety. However, the challenge of assuri…

Alignment Plausibility: A New Standard for Assuring AI in Healthcare

2026-07-08 · Gwydion Williams, Sara Zannone, Bilal A Mateen arxiv

Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the fric…