paper-with-me

홈 › Papers

Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements

2023-02-18 · Jiawen Deng, Jiale Cheng, Hao Sun, Zhexin Zhang, Minlie Huang

As generative large model capabilities advance, safety concerns become more pronounced in their outputs. To ensure the sustainable growth of the AI ecosystem, it's imperative to undertake a holistic evaluation and refinement of associated safety risks. This survey presents a framework for safety research pertaining to large models, delineating the landscape of safety risks as well as safety evaluation and improvement methods. We begin by introducing safety issues of wide concern, then delve into safety evaluation methods for large models, encompassing preference-based testing, adversarial attack approaches, issues detection, and other advanced evaluation methods. Additionally, we explore the strategies for enhancing large model safety from training to deployment, highlighting cutting-edge safety approaches for each stage in building large models. Finally, we discuss the core challenges in advancing towards more responsible AI, including the interpretability of safety mechanisms, ongoing safety issues, and robustness against malicious attacks. Through this survey, we aim to provide clear technical guidance for safety researchers and encourage further study on the safety of large models.

📄 PDF Abstract BibTeX arXiv:2302.09270

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial AttackEthicsSurvey

Similar Papers 제목 키워드 기반

Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data

2025-05-15 · Adel ElZemity, Budi Arief, Shujun Li

The integration of large language models (LLMs) into cyber security applications presents significant opportunities, such as enhancing threat analysis and malware detection, but can also introduce critical risks and safe…

Malware DetectionSafety Alignment

AI Safety in Generative AI Large Language Models: A Survey

2024-07-06 · Jaymari Chua, Yun Li, Shiyi Yang, Chen Wang 외

Large Language Model (LLMs) such as ChatGPT that exhibit generative AI capabilities are facing accelerated adoption and innovation. The increased presence of Generative AI (GAI) inevitably raises concerns about the risks…

Language ModellingLarge Language ModelSurvey

SafeRun: Enabling Determinism in LLM Planning for Running

2026-06-08 · Meilin Chen, Zepeng Zhai, Jiaxuan Zhao, Yuan Lu arxiv

Large Language Models enable flexible natural-language planning but remain unreliable in determinism-critical domains due to their probabilistic nature. This limitation is especially problematic in running planning, wher…

MolSafeEval: A Benchmark for Uncovering Safety Risks in AI-Generated Molecules

2026-07-01 · Tong Xu, Xinzhe Cao, Zhihui Zhu, Keyan Ding 외 arxiv

Current molecular generation benchmarks emphasize task complexity, molecule novelty, and property alignment; they largely overlook a critical concern: the potential safety risks of AI-generated molecules. In practice, ma…

SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models

2025-10-08 · Huahui Yi, Kun Wang, Qiankun Li, Miao Yu 외 arxiv

Multimodal Large Reasoning Models (MLRMs) demonstrate impressive cross-modal reasoning but often amplify safety risks under adversarial or unsafe prompts, a phenomenon we call the \textit{Reasoning Tax}. Existing defense…

Reinforcement LearningMultimodal Reasoning