paper-with-me

Papers

KZ-SafetyPrompts: A Kazakh Safety Evaluation Prompt Dataset for Large Language Models

2026-05-26 · Wajdi Zaghouani, Shimaa Amer Ibrahim, Aruzhan Muratbek, Olzhasbek Zhakenov, Adiya Akhmetzhanova arxiv

Kazakh is underrepresented in resources for evaluating the safety behavior of large language models. We present KZ-SafetyPrompts, a Kazakh prompt dataset for safety evaluation across eleven categories covering common risk areas such as self-harm, violence, child exploitation, sexual content, racist content, radicalization, and regulated goods or illegal activities. The dataset contains 5,717 prompts written natively in Kazakh (Cyrillic), organized by category, with English translations for cross-lingual analysis. Prompts resemble realistic user queries, often in a teen or child style, and are phrased as intent prompts without procedural instructions. We document the writing protocol, labeling procedures (including borderline-case decision rules), and quality-control steps (schema standardization, completeness checks, and deduplication). We also align the categories with widely used safety taxonomies to support integration with existing evaluation pipelines. Baseline results with GPT-4o show an overall refusal rate of 28.2%, varying from 5.5% to 53.8% across categories, indicating that Kazakh prompts expose category-specific safety gaps not captured by English-only evaluation.

📄 PDF Abstract BibTeX arXiv:2605.26947

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety

2024-04-08 · Paul Röttger, Fabio Pernisi, Bertie Vidgen, Dirk Hovy

The last two years have seen a rapid growth in concerns around the safety of large language models (LLMs). Researchers and practitioners have met these concerns by creating an abundance of datasets for evaluating and imp…

Language ModelingLanguage ModellingLarge Language Model

Qorgau: Evaluating LLM Safety in Kazakh-Russian Bilingual Contexts

2025-02-19 · Maiya Goloburda, Nurkhan Laiyk, Diana Turmakhan, Yuxia Wang 외

Large language models (LLMs) are known to have the potential to generate harmful content, posing risks to users. While significant progress has been made in developing taxonomies for LLM risks and safety evaluation promp…

Safety Assessment of Chinese Large Language Models

2023-04-20 · Hao Sun, Zhexin Zhang, Jiawen Deng, Jiale Cheng 외

With the rapid popularity of large language models such as ChatGPT and GPT-4, a growing amount of attention is paid to their safety concerns. These models may generate insulting and discriminatory content, reflect incorr…

Llama-3.1-Sherkala-8B-Chat: An Open Large Language Model for Kazakh

2025-03-03 · Fajri Koto, Rituraj Joshi, Nurdaulet Mukhituly, Yuxia Wang 외

Llama-3.1-Sherkala-8B-Chat, or Sherkala-Chat (8B) for short, is a state-of-the-art instruction-tuned open generative large language model (LLM) designed for Kazakh. Sherkala-Chat (8B) aims to enhance the inclusivity of L…

Language ModelingLanguage ModellingLarge Language ModelSafety Alignment

No One-Size-Fits-All: Building Systems For Translation to Bashkir, Kazakh, Kyrgyz, Tatar and Chuvash Using Synthetic And Original Data

2026-02-04 · Dmitry Karpov arxiv

We explore machine translation for five Turkic language pairs: Russian-Bashkir, Russian-Kazakh, Russian-Kyrgyz, English-Tatar, English-Chuvash. Fine-tuning nllb-200-distilled-600M with LoRA on synthetic data achieved chr…

Machine Translation