paper-with-me

Papers

Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings

2025-03-19 · Zonghao Ying, Guangyi Zheng, Yongxin Huang, Deyue Zhang, Wenxin Zhang, Quanchen Zou, Aishan Liu, Xianglong Liu, DaCheng Tao

This study presents the first comprehensive safety evaluation of the DeepSeek models, focusing on evaluating the safety risks associated with their generated content. Our evaluation encompasses DeepSeek's latest generation of large language models, multimodal large language models, and text-to-image models, systematically examining their performance regarding unsafe content generation. Notably, we developed a bilingual (Chinese-English) safety evaluation dataset tailored to Chinese sociocultural contexts, enabling a more thorough evaluation of the safety capabilities of Chinese-developed models. Experimental results indicate that despite their strong general capabilities, DeepSeek models exhibit significant safety vulnerabilities across multiple risk dimensions, including algorithmic discrimination and sexual content. These findings provide crucial insights for understanding and improving the safety of large foundation models. Our code is available at https://github.com/NY1024/DeepSeek-Safety-Eval.

📄 PDF Abstract BibTeX arXiv:2503.15092

Code (1)

ny1024/deepseek-safety-eval 공식 구현

Similar Papers 제목 키워드 기반

Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts

2025-03-18 · Wenjing Zhang, Xuejiao Lei, Zhaoxiang Liu, Limin Han 외

DeepSeek-R1, renowned for its exceptional reasoning capabilities and open-source strategy, is significantly influencing the global artificial intelligence landscape. However, it exhibits notable safety shortcomings. Rece…

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations

2025-02-14 · Mang Ye, Xuankun Rong, Wenke Huang, Bo Du 외

With the rapid advancement of Large Vision-Language Models (LVLMs), ensuring their safety has emerged as a crucial area of research. This survey provides a comprehensive analysis of LVLM safety, covering key aspects such…

Survey

Safety Evaluation of DeepSeek Models in Chinese Contexts

2025-02-16 · Wenjing Zhang, Xuejiao Lei, Zhaoxiang Liu, Ning Wang 외

Recently, the DeepSeek series of models, leveraging their exceptional reasoning capabilities and open-source strategy, is reshaping the global AI landscape. Despite these advantages, they exhibit significant safety defic…

DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning

2025-04-02 · Sara Vera Marjanović, Arkil Patel, Vaibhav Adlakha, Milad Aghajohari 외

Large Reasoning Models like DeepSeek-R1 mark a fundamental shift in how LLMs approach complex problems. Instead of directly producing an answer for a given input, DeepSeek-R1 creates detailed multi-step reasoning chains,…

Discovering Forbidden Topics in Language Models

2025-05-23 · Can Rager, Chris Wendler, Rohit Gandikota, David Bau

Refusal discovery is the task of identifying the full set of topics that a language model refuses to discuss. We introduce this new problem setting and develop a refusal discovery method, LLM-crawler, that uses token pre…

Memorization