paper-with-me

홈 › Papers

Defining and Evaluating Physical Safety for Large Language Models

2024-11-04 · Yung-Chen Tang, Pin-Yu Chen, Tsung-Yi Ho

Large Language Models (LLMs) are increasingly used to control robotic systems such as drones, but their risks of causing physical threats and harm in real-world applications remain unexplored. Our study addresses the critical gap in evaluating LLM physical safety by developing a comprehensive benchmark for drone control. We classify the physical safety risks of drones into four categories: (1) human-targeted threats, (2) object-targeted threats, (3) infrastructure attacks, and (4) regulatory violations. Our evaluation of mainstream LLMs reveals an undesirable trade-off between utility and safety, with models that excel in code generation often performing poorly in crucial safety aspects. Furthermore, while incorporating advanced prompt engineering techniques such as In-Context Learning and Chain-of-Thought can improve safety, these methods still struggle to identify unintentional attacks. In addition, larger models demonstrate better safety capabilities, particularly in refusing dangerous commands. Our findings and benchmark can facilitate the design and evaluation of physical safety for LLMs. The project page is available at huggingface.co/spaces/TrustSafeAI/LLM-physical-safety.

📄 PDF Abstract BibTeX arXiv:2411.02317

Code (0)

등록된 구현이 없습니다.

Tasks

Code GenerationIn-Context LearningPrompt Engineering

Similar Papers 제목 키워드 기반

Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models

2024-09-20 · Hao Cheng, Erjia Xiao, Chengyuan Yu, Zhao Yao 외

Recently, driven by advancements in Multimodal Large Language Models (MLLMs), Vision Language Action Models (VLAMs) are being proposed to achieve better performance in open-vocabulary scenarios for robotic manipulation t…

Vision-Language-Action

Toward Reliable, Safe, and Secure LLMs for Scientific Applications

2026-03-18 · Saket Sanjeev Chaturvedi, Joshua Bergerson, Tanwi Mallick arxiv

As large language models (LLMs) evolve into autonomous "AI scientists," they promise transformative advances but introduce novel vulnerabilities, from potential "biosafety risks" to "dangerous explosions." Ensuring trust…

SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents

2025-10-14 · Simon Sinong Zhan, Yao Liu, Philip Wang, Zinan Wang 외 arxiv

We present SENTINEL, a framework for formally evaluating the physical safety of foundation model (FM)-based embodied agents. SENTINEL is the first to provide multi-level safety evaluation across semantic interpretation, …

Evaluating Reinforcement Learning Safety and Trustworthiness in Cyber-Physical Systems

2025-03-12 · Katherine Dearstyne, Pedro, Alarcon Granadeno, Theodore Chambers 외

Cyber-Physical Systems (CPS) often leverage Reinforcement Learning (RL) techniques to adapt dynamically to changing environments and optimize performance. However, it is challenging to construct safety cases for RL compo…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models

2026-04-21 · Josue Torres-Fonseca, Naihao Deng, Yinpei Dai, Shane Storks 외 arxiv

Multimodal Large Language Models are increasingly adopted as autonomous agents in interactive environments, yet their ability to proactively address safety hazards remains insufficient. We introduce SafetyALFRED, built u…

Question Answering