paper-with-me

홈 › Papers

The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts

2024-01-23 · Lingfeng Shen, Weiting Tan, Sihao Chen, Yunmo Chen, Jingyu Zhang, Haoran Xu, Boyuan Zheng, Philipp Koehn, Daniel Khashabi

As the influence of large language models (LLMs) spans across global communities, their safety challenges in multilingual settings become paramount for alignment research. This paper examines the variations in safety challenges faced by LLMs across different languages and discusses approaches to alleviating such concerns. By comparing how state-of-the-art LLMs respond to the same set of malicious prompts written in higher- vs. lower-resource languages, we observe that (1) LLMs tend to generate unsafe responses much more often when a malicious prompt is written in a lower-resource language, and (2) LLMs tend to generate more irrelevant responses to malicious prompts in lower-resource languages. To understand where the discrepancy can be attributed, we study the effect of instruction tuning with reinforcement learning from human feedback (RLHF) or supervised finetuning (SFT) on the HH-RLHF dataset. Surprisingly, while training with high-resource languages improves model alignment, training in lower-resource languages yields minimal improvement. This suggests that the bottleneck of cross-lingual alignment is rooted in the pretraining stage. Our findings highlight the challenges in cross-lingual LLM safety, and we hope they inform future research in this direction.

📄 PDF Abstract BibTeX arXiv:2401.13136

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

BarrierSteer: LLM Safety via Learning Barrier Steering

2026-02-23 · Thanh Q. Tran, Arun Verma, Kiwan Wong, Bryan Kian Hsiang Low 외 arxiv

Despite the strong performance of large language models (LLMs) across diverse tasks, their susceptibility to adversarial attacks and unsafe content generation remains a significant obstacle to deployment, particularly in…

Adversarial Attack

Control Barrier Function for Aligning Large Language Models

2025-11-05 · Yuya Miyaoka, Masaki Inoue arxiv

This paper proposes a control-based framework for aligning large language models (LLMs) by leveraging a control barrier function (CBF) to ensure user-desirable text generation. The presented framework applies the CBF saf…

Text Generation

The Multilingual Divide and Its Impact on Global AI Safety

2025-05-27 · Aidan Peppin, Julia Kreutzer, Alice Schoenauer Sebag, Kelly Marchisio 외

Despite advances in large language model capabilities in recent years, a large gap remains in their capabilities and safety performance for many languages beyond a relatively small handful of globally dominant languages.…

Language ModelingLanguage ModellingLarge Language Model

Investigating the Impact of Quantization Methods on the Safety and Reliability of Large Language Models

2025-02-18 · Artyom Kharinaev, Viktor Moskvoretskii, Egor Shvetsov, Kseniia Studenikina 외

Large Language Models (LLMs) have emerged as powerful tools for addressing modern challenges and enabling practical applications. However, their computational expense remains a significant barrier to widespread adoption.…

Quantization

A Lightweight Multi-Agent Framework for Automated Concrete Barrier Design

2026-06-10 · Wanting Wang, Xiye Ma, Yuyang He, Minghui Cheng 외 arxiv

The design of reinforced concrete highway barriers is a safety-critical process that requires strict compliance with regulatory provisions such as the AASHTO-LRFD bridge design guidelines. Current engineering practice re…