paper-with-me

Papers

AlbanianLLMSafety: A Safety Evaluation Dataset for Large Language Models in Albanian

2026-05-26 · Wajdi Zaghouani, Kholoud K. Aldous, Isra Fejzullaj arxiv

Safety evaluation of Large Language Models (LLMs) has largely focused on high-resource languages, leaving low-resource languages critically underserved. We present AlbanianLLMSafety, the first publicly available safety evaluation dataset for LLMs in Albanian, a linguistically distinct low-resource language with approximately 7.5 million speakers across Albania, Kosovo, North Macedonia, and the diaspora. The dataset contains 2,951 prompts spanning 11 safety categories, including self-harm, violence, racist content, child exploitation, and radicalization, with an average of 268 prompts per category. Each prompt is provided in Albanian with an English reference translation and a detailed category label. This resource addresses a significant gap in safety evaluation infrastruc-ture for low-resource languages and provides an essential benchmark for developing safer, more inclusive LLMs. The dataset will be provided upon request to support safety evaluation, fine-tuning, red-teaming, and guardrail development for Albanian-speaking communities.

📄 PDF Abstract BibTeX arXiv:2605.26954

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LinguaSafe: A Comprehensive Multilingual Safety Benchmark for Large Language Models

2025-08-18 · Zhiyuan Ning, Tianle Gu, Jiaxin Song, Shixin Hong 외 arxiv

The widespread adoption and increasing prominence of large language models (LLMs) in global technologies necessitate a rigorous focus on ensuring their safety across a diverse range of linguistic and cultural contexts. T…

Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings

2025-03-19 · Zonghao Ying, Guangyi Zheng, Yongxin Huang, Deyue Zhang 외

This study presents the first comprehensive safety evaluation of the DeepSeek models, focusing on evaluating the safety risks associated with their generated content. Our evaluation encompasses DeepSeek's latest generati…

Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models

2025-06-25 · KeJia Chen, Jiawen Zhang, Jiacong Hu, Yu Wang 외

Quantized large language models (LLMs) have gained increasing attention and significance for enabling deployment in resource-constrained environments. However, emerging studies on a few calibration dataset-free quantizat…

Quantization

RabakBench: Scaling Human Annotations to Construct Localized Multilingual Safety Benchmarks for Low-Resource Languages

2025-07-08 · Gabriel Chua, Leanne Tan, Ziyu Ge, Roy Ka-Wei Lee

Large language models (LLMs) and their safety classifiers often perform poorly on low-resource languages due to limited training data and evaluation benchmarks. This paper introduces RabakBench, a new multilingual safety…

Red Teaming

KZ-SafetyPrompts: A Kazakh Safety Evaluation Prompt Dataset for Large Language Models

2026-05-26 · Wajdi Zaghouani, Shimaa Amer Ibrahim, Aruzhan Muratbek, Olzhasbek Zhakenov 외 arxiv

Kazakh is underrepresented in resources for evaluating the safety behavior of large language models. We present KZ-SafetyPrompts, a Kazakh prompt dataset for safety evaluation across eleven categories covering common ris…