paper-with-me

홈 › Papers

SAIF: A Comprehensive Framework for Evaluating the Risks of Generative AI in the Public Sector

2025-01-15 · Kyeongryul Lee, Heehyeon Kim, Joyce Jiyoung Whang

The rapid adoption of generative AI in the public sector, encompassing diverse applications ranging from automated public assistance to welfare services and immigration processes, highlights its transformative potential while underscoring the pressing need for thorough risk assessments. Despite its growing presence, evaluations of risks associated with AI-driven systems in the public sector remain insufficiently explored. Building upon an established taxonomy of AI risks derived from diverse government policies and corporate guidelines, we investigate the critical risks posed by generative AI in the public sector while extending the scope to account for its multimodal capabilities. In addition, we propose a Systematic dAta generatIon Framework for evaluating the risks of generative AI (SAIF). SAIF involves four key stages: breaking down risks, designing scenarios, applying jailbreak methods, and exploring prompt types. It ensures the systematic and consistent generation of prompt data, facilitating a comprehensive evaluation while providing a solid foundation for mitigating the risks. Furthermore, SAIF is designed to accommodate emerging jailbreak methods and evolving prompt types, thereby enabling effective responses to unforeseen risk scenarios. We believe that this study can play a crucial role in fostering the safe and responsible integration of generative AI into the public sector.

📄 PDF Abstract BibTeX arXiv:2501.08814

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RAIL: A modular framework for Reinforcement-learning-based Adversarial Imitation Learning

2021-05-08 · Eddy Hudson, Garrett Warnell, Peter Stone

While Adversarial Imitation Learning (AIL) algorithms have recently led to state-of-the-art results on various imitation learning benchmarks, it is unclear as to what impact various design decisions have on performance. …

Imitation LearningOpenAI Gymreinforcement-learningReinforcement Learning (RL)

A Formal Framework for Assessing and Mitigating Emergent Security Risks in Generative AI Models: Bridging Theory and Dynamic Risk Mitigation

2024-10-15 · Aviral Srivastava, Sourav Panda

As generative AI systems, including large language models (LLMs) and diffusion models, advance rapidly, their growing adoption has led to new and complex security risks often overlooked in traditional AI risk assessment …

Anomaly DetectionRed Teaming

SAIF: A Stability-Aware Inference Framework for Medical Image Segmentation with Segment Anything Model

2026-03-13 · Ke Wu, Shiqi Chen, Yiheng Zhong, Hengxian Liu 외 arxiv

Segment Anything Model (SAM) enable scalable medical image segmentation but suffer from inference-time instability when deployed as a frozen backbone. In practice, bounding-box prompts often contain localization errors, …

Medical Image Segmentation

Sociotechnical Safety Evaluation of Generative AI Systems

2023-10-18 · Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arianna Manzini 외

Generative AI systems produce a range of risks. To ensure the safety of generative AI systems, these risks must be evaluated. In this paper, we make two main contributions toward establishing such evaluations. First, we …

Safe Active Feature Selection for Sparse Learning

2018-06-15 · Shaogang Ren, Jianhua Z. Huang, Shuai Huang, Xiaoning Qian

We present safe active incremental feature selection~(SAIF) to scale up the computation of LASSO solutions. SAIF does not require a solution from a heavier penalty parameter as in sequential screening or updating the ful…

feature selectionSparse Learning