paper-with-me

홈 › Papers

SomaliBench Eval: Measuring English-to-Somali Refusal Gaps in Open-Weight Language Models

2026-05-25 · Khalid Yusuf Dahir arxiv

Large language model safety evaluation remains heavily English-centered, leaving low-resource languages under-measured even when models are deployed globally. We evaluate four open-weight instruction-tuned models on SomaliBench v0, a native-author-verified benchmark of 100 harmful-intent prompts paired across English and Somali. Each of Llama-3.1-8B-Instruct, Gemma-2-9B-Instruct, Qwen-2.5-7B-Instruct, and Aya-23-8B is run locally with temperature 0 and the same English "helpful, harmless, and honest" (HHH) system prompt. We find large English-to-Somali refusal gaps for all four models, ranging from 0.40 to 0.93, all strictly positive under a paired bootstrap and significant by exact McNemar tests. For three models, the dominant Somali non-refusal mode is not fluent harmful compliance but unclear output: wrong-language, incoherent, or off-topic generations. A pinned Claude Sonnet snapshot (claude-sonnet-4-5-20250929) classifies each response as refused, complied, or unclear; its safety layer declined 34 of 800 classifications, which the native author labeled manually. A native-author spot-check achieves 100% agreement with the judge (Cohen's $κ=1.00$) on 74 comparable rows. We report aggregate refusal rates, category gaps, and reliability statistics only; raw model generations are retained locally and are not released because some may contain harmful content.

📄 PDF Abstract BibTeX arXiv:2605.25420

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LSR: Linguistic Safety Robustness Benchmark for Low-Resource West African Languages

2026-02-27 · Godwin Abuh Faruna arxiv

Safety alignment in large language models relies predominantly on English-language training data. When harmful intent is expressed in low-resource languages, refusal mechanisms that hold in English frequently fail to act…

Evaluating Multilingual Long-Context Models for Retrieval and Reasoning

2024-09-26 · Ameeta Agrawal, Andy Dang, Sina Bagheri Nezhad, Rhitabrat Pokharel 외

Recent large language models (LLMs) demonstrate impressive capabilities in handling long contexts, some exhibiting near-perfect recall on synthetic retrieval tasks. However, these evaluations have mainly focused on Engli…

RetrievalSentence

Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts

2026-06-22 · Arthur Wuhrmann, Gaetan Stein, Daniel Brunner, Andrei Kucharavy arxiv

While the wider applicability of LLMs in the legal field is currently debated due to their reliability and the gravity of any errors, narrow uses with well-understood and mitigated risks have emerged. Notably the Swiss F…

SEARCHER: Shared Embedding Architecture for Effective Retrieval

2020-05-01 · LREC 2020 5 · Joel Barry, Elizabeth Boschee, Marjorie Freedman, Scott Miller

We describe an approach to cross lingual information retrieval that does not rely on explicit translation of either document or query terms. Instead, both queries and documents are mapped into a shared embedding space wh…

Cross-Lingual Information RetrievalInformation RetrievalRetrievalTranslation

SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark

2026-05-18 · Khalid Yusuf Dahir arxiv

Somali is a Cushitic language of the Horn of Africa with ~25 million speakers, yet no documented dedicated Somali pretraining corpus with a companion tokenizer and language-identification benchmark has been publicly rele…