paper-with-me

홈 › Papers

Qorgau: Evaluating LLM Safety in Kazakh-Russian Bilingual Contexts

2025-02-19 · Maiya Goloburda, Nurkhan Laiyk, Diana Turmakhan, Yuxia Wang, Mukhammed Togmanov, Jonibek Mansurov, Askhat Sametov, Nurdaulet Mukhituly, Minghan Wang, Daniil Orel, Zain Muhammad Mujahid, Fajri Koto, Timothy Baldwin, Preslav Nakov

Large language models (LLMs) are known to have the potential to generate harmful content, posing risks to users. While significant progress has been made in developing taxonomies for LLM risks and safety evaluation prompts, most studies have focused on monolingual contexts, primarily in English. However, language- and region-specific risks in bilingual contexts are often overlooked, and core findings can diverge from those in monolingual settings. In this paper, we introduce Qorgau, a novel dataset specifically designed for safety evaluation in Kazakh and Russian, reflecting the unique bilingual context in Kazakhstan, where both Kazakh (a low-resource language) and Russian (a high-resource language) are spoken. Experiments with both multilingual and language-specific LLMs reveal notable differences in safety performance, emphasizing the need for tailored, region-specific datasets to ensure the responsible and safe deployment of LLMs in countries like Kazakhstan. Warning: this paper contains example data that may be offensive, harmful, or biased.

📄 PDF Abstract BibTeX arXiv:2502.13640

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Llama-3.1-Sherkala-8B-Chat: An Open Large Language Model for Kazakh

2025-03-03 · Fajri Koto, Rituraj Joshi, Nurdaulet Mukhituly, Yuxia Wang 외

Llama-3.1-Sherkala-8B-Chat, or Sherkala-Chat (8B) for short, is a state-of-the-art instruction-tuned open generative large language model (LLM) designed for Kazakh. Sherkala-Chat (8B) aims to enhance the inclusivity of L…

Language ModelingLanguage ModellingLarge Language ModelSafety Alignment

The TALP-UPC Machine Translation Systems for WMT19 News Translation Task: Pivoting Techniques for Low Resource MT

2019-08-01 · WS 2019 8 · Noe Casas, Jos{\'e} A. R. Fonollosa, Carlos Escolano, Christine Basta 외

In this article, we describe the TALP-UPC research group participation in the WMT19 news translation shared task for Kazakh-English. Given the low amount of parallel training data, we resort to using Russian as pivot lan…

DecoderMachine TranslationTranslation

Loanword or Switch? The Annotation Boundary, Not the Model, Drives Kazakh-Russian Code-Switching Identification

2026-08-01 · Bogdan Savelyev arxiv

Off-the-shelf LID and letter heuristics over-label Kazakh-Russian social text as mixed: Russian loanwords inside Kazakh look like code-switching under a shared Cyrillic script. We release a document-level gold LID set wh…

Multi-Source Transformer for Kazakh-Russian-English Neural Machine Translation

2019-08-01 · WS 2019 8 · Patrick Littell, Chi-kiu Lo, Samuel Larkin, Darlene Stewart

We describe the neural machine translation (NMT) system developed at the National Research Council of Canada (NRC) for the Kazakh-English news translation task of the Fourth Conference on Machine Translation (WMT19). Our…

Machine TranslationNMTSentenceTranslation

Low-resource Machine Translation for Code-switched Kazakh-Russian Language Pair

2025-03-25 · Maksim Borisov, Zhanibek Kozhirbayev, Valentin Malykh

Machine translation for low resource language pairs is a challenging task. This task could become extremely difficult once a speaker uses code switching. We propose a method to build a machine translation model for code-…

Machine TranslationTranslation