paper-with-me

홈 › Papers

An Independent Safety Evaluation of Kimi K2.5

2026-04-03 · Zheng-Xin Yong, Parv Mahajan, Andy Wang, Ida Caspary, Yernat Yestekov, Zora Che, Mosh Levy, Elle Najt, Dennis Murphy, Prashant Kulkarni, Lev McKinney, Kei Nishimura-Gasparian, Ram Potham, Aengus Lynch, Michael L. Chen arxiv

Kimi K2.5 is an open-weight LLM that rivals closed models across coding, multimodal, and agentic benchmarks, but was released without an accompanying safety evaluation. In this work, we conduct a preliminary safety assessment of Kimi K2.5 focusing on risks likely to be exacerbated by powerful open-weight models. Specifically, we evaluate the model for CBRNE misuse risk, cybersecurity risk, misalignment, political censorship, bias, and harmlessness, in both agentic and non-agentic settings. We find that Kimi K2.5 shows similar dual-use capabilities to GPT 5.2 and Claude Opus 4.5, but with significantly fewer refusals on CBRNE-related requests, suggesting it may uplift malicious actors in weapon creation. On cyber-related tasks, we find that Kimi K2.5 demonstrates competitive cybersecurity performance, but it does not appear to possess frontier-level autonomous cyberoffensive capabilities such as vulnerability discovery and exploitation. We further find that Kimi K2.5 shows concerning levels of sabotage ability and self-replication propensity, although it does not appear to have long-term malicious goals. In addition, Kimi K2.5 exhibits narrow censorship and political bias, especially in Chinese, and is more compliant with harmful requests related to spreading disinformation and copyright infringement. Finally, we find the model refuses to engage in user delusions and generally has low over-refusal rates. While preliminary, our findings highlight how safety risks exist in frontier open-weight models and may be amplified by the scale and accessibility of open-weight releases. Therefore, we strongly urge open-weight model developers to conduct and release more systematic safety evaluations required for responsible deployment.

📄 PDF Abstract BibTeX arXiv:2604.03121

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Kimi-Audio Technical Report

2025-04-25 · KimiTeam, Ding Ding, Zeqian Ju, Yichong Leng 외

We present Kimi-Audio, an open-source audio foundation model that excels in audio understanding, generation, and conversation. We detail the practices in building Kimi-Audio, including model architecture, data curation, …

Audio Question AnsweringQuestion Answeringspeech-recognitionSpeech Recognition

Kimi K3: Open Frontier Intelligence

2026-07-27 · Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao 외 hf

We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attent…

Reinforcement Learning

Kimi-VL Technical Report

2025-04-10 · Kimi Team, Angang Du, Bohong Yin, Bowei Xing 외

We present Kimi-VL, an efficient open-source Mixture-of-Experts (MoE) vision-language model (VLM) that offers advanced multimodal reasoning, long-context understanding, and strong agent capabilities - all while activatin…

Long-Context UnderstandingMathematical ReasoningMixture-of-ExpertsMultimodal Reasoning+2

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis

2026-03-20 · May Lynn Reese, Markela Zeneli, Mindy Ng, Jacob Haimes 외 arxiv

General-purpose Large Language Models (LLMs) are becoming widely adopted by people for mental health support. Yet emerging evidence suggests there are significant risks associated with high-frequency use, particularly fo…

Kimi K2.5: Visual Agentic Intelligence

2026-02-02 · Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao 외 arxiv

We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. This …

Reinforcement Learning