paper-with-me

Papers

ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI

2026-02-15 · Haibo Tong, Feifei Zhao, Linghao Feng, Ruoyu Wu, Ruolin Chen, Lu Jia, Zhou Zhao, Jindong Li, Tenglong Li, Erliang Lin, Shuai Yang, Enmeng Lu, Yinqian Sun, Qian Zhang, Zizhe Ruan, Jinyu Fan, Zeyang Yue, Ping Wu, Huangrui Li, Chengyi Sun, Yi Zeng arxiv

Rapidly evolving AI exhibits increasingly strong autonomy and goal-directed capabilities, accompanied by derivative systemic risks that are more unpredictable, difficult to control, and potentially irreversible. However, current AI safety evaluation systems suffer from critical limitations such as restricted risk dimensions and failed frontier risk detection. The lagging safety benchmarks and alignment technologies can hardly address the complex challenges posed by cutting-edge AI models. To bridge this gap, we propose the "ForesightSafety Bench" AI Safety Evaluation Framework, beginning with 7 major Fundamental Safety pillars and progressively extends to advanced Embodied AI Safety, AI4Science Safety, Social and Environmental AI risks, Catastrophic and Existential Risks, as well as 8 critical industrial safety domains, forming a total of 94 refined risk dimensions. To date, the benchmark has accumulated tens of thousands of structured risk data points and assessment results, establishing a widely encompassing, hierarchically clear, and dynamically evolving AI safety evaluation framework. Based on this benchmark, we conduct systematic evaluation and in-depth analysis of over twenty mainstream advanced large models, identifying key risk patterns and their capability boundaries. The safety capability evaluation results reveals the widespread safety vulnerabilities of frontier AI across multiple pillars, particularly focusing on Risky Agentic Autonomy, AI4Science Safety, Embodied AI Safety, Social AI Safety and Catastrophic and Existential Risks. Our benchmark is released at https://github.com/Beijing-AISI/ForesightSafety-Bench. The project website is available at https://foresightsafety-bench.beijing-aisi.ac.cn/.

📄 PDF Abstract BibTeX arXiv:2602.14135

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models

2026-06-25 · Mingyang Lyu, Yinqian Sun, Yiyang Jia, Sicheng Shen 외 arxiv

In embodied intelligence, safety is a prerequisite for reliable robot deployment in the physical world. Current vision-language-action (VLA) models continue to advance toward general-purpose task capability, yet their em…

An International Consortium for Evaluations of Societal-Scale Risks from Advanced AI

2023-10-22 · Ross Gruetzemacher, Alan Chan, Kevin Frazier, Christy Manning 외

Given rapid progress toward advanced AI and risks from frontier AI systems (advanced AI systems pushing the boundaries of the AI capabilities frontier), the creation and implementation of AI governance and regulatory sch…

Deprecating Benchmarks: Criteria and Framework

2025-07-08 · Ayrton San Joaquin, Rokas Gipiškis, Leon Staufer, Ariel Gil arxiv

As frontier artificial intelligence (AI) models rapidly advance, benchmarks are integral to comparing different models and measuring their progress in different task-specific domains. However, there is a lack of guidance…

Data-Centric AI Governance: Addressing the Limitations of Model-Focused Policies

2024-09-25 · Ritwik Gupta, Leah Walker, Rodolfo Corona, Stephanie Fu 외

Current regulations on powerful AI capabilities are narrowly focused on "foundation" or "frontier" models. However, these terms are vague and inconsistently defined, leading to an unstable foundation for governance effor…

Opportunities and Challenges of Frontier Data Governance With Synthetic Data

2025-03-21 · Madhavendra Thakur, Jason Hausenloy

Synthetic data, or data generated by machine learning models, is increasingly emerging as a solution to the data access problem. However, its use introduces significant governance and accountability challenges, and poten…