Foundational Challenges in Assuring Alignment and Safety of Large Language Models
This work identifies 18 foundational challenges in assuring the alignment and safety of large language models (LLMs). These challenges are organized into three different categories: scientific understanding of LLMs, development and deployment methods, and sociotechnical challenges. Based on the identified challenges, we pose $200+$ concrete research questions.
Code (1)
Similar Papers 제목 키워드 기반
Assuring the Safety of Reinforcement Learning Components: AMLAS-RL
The rapid advancement of machine learning (ML) has led to its increasing integration into cyber-physical systems (CPS) across diverse domains. While CPS offer powerful capabilities, incorporating ML components introduces…
Reinforcement LearningSafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning
Vision-language-action models (VLAs) show potential as generalist robot policies. However, these models pose extreme safety challenges during real-world deployment, including the risk of harm to the environment, the robo…
Safe Reinforcement LearningSafety AlignmentVision-Language-ActionTesting and verification of neural-network-based safety-critical control software: A systematic literature review
Context: Neural Network (NN) algorithms have been successfully adopted in a number of Safety-Critical Cyber-Physical Systems (SCCPSs). Testing and Verification (T&V) of NN-based control software in safety-critical domain…
Systematic Literature ReviewLandscape of AI safety concerns -- A methodology to support safety assurance for AI-based autonomous systems
Artificial Intelligence (AI) has emerged as a key technology, driving advancements across a range of applications. Its integration into modern autonomous systems requires assuring safety. However, the challenge of assuri…
Alignment Plausibility: A New Standard for Assuring AI in Healthcare
Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the fric…