paper-with-me

홈 › Papers

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense

2024-12-31 · Keke Zhai

Currently, large models are prone to generating harmful content when faced with complex attack instructions, significantly reducing their defensive capabilities. To address this issue, this paper proposes a method based on constructing data aligned with multi-dimensional attack defense to enhance the generative security of large models. The core of our method lies in improving the effectiveness of safe alignment learning for large models by innova-tively increasing the diversity of attack instruction dimensions and the accuracy of generat-ing safe responses. To validate the effectiveness of our method, beyond existing security evaluation benchmarks, we additionally designed new security evaluation benchmarks and conducted comparative experiments using Llama3.2 as the baseline model. The final ex-perimental results demonstrate that our method can significantly improve the generative security of large models under complex instructional attacks, while also maintaining and enhancing the models' general capabilities.

📄 PDF Abstract BibTeX arXiv:2501.00517

Code (0)

등록된 구현이 없습니다.

Tasks

Diversity

Similar Papers 제목 키워드 기반

MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving

2025-12-04 · Bin Sun, Yaoguang Cao, Yan Wang, Rui Wang 외 arxiv

End-to-End autonomous driving (E2E-AD) has emerged as a new paradigm, where trajectory planning plays a crucial role. Existing studies mainly follow two directions: trajectory generation oriented, which focuses on produc…

Trajectory PlanningAutonomous DrivingDecision Making

Enhancing the Performance of DeepReach on High-Dimensional Systems through Optimizing Activation Functions

2023-12-29 · Qian Wang, Tianhao Wu

With the continuous advancement in autonomous systems, it becomes crucial to provide robust safety guarantees for safety-critical systems. Hamilton-Jacobi Reachability Analysis is a formal verification method that guaran…

SLM as Guardian: Pioneering AI Safety with Small Language Models

2024-05-30 · Ohjoon Kwon, Donghyeon Jeon, Nayoung Choi, Gyu-Hwung Cho 외

Most prior safety research of large language models (LLMs) has focused on enhancing the alignment of LLMs to better suit the safety requirements of humans. However, internalizing such safeguard features into larger model…

Multi-Task LearningResponse Generation

Model-free Neural Lyapunov Control for Safe Robot Navigation

2022-03-02 · Zikang Xiong, Joe Eappen, Ahmed H. Qureshi, Suresh Jagannathan

Model-free Deep Reinforcement Learning (DRL) controllers have demonstrated promising results on various challenging non-linear control tasks. While a model-free DRL algorithm can solve unknown dynamics and high-dimension…

Deep Reinforcement LearningRobot Navigation

Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models

2025-06-02 · Youze Wang, WenBo Hu, Yinpeng Dong, Jing Liu 외

Large Language Models (LLMs) have evolved into Multimodal Large Language Models (MLLMs), significantly enhancing their capabilities by integrating visual information and other types, thus aligning more closely with the n…

Safety Alignment