paper-with-me

Papers

This Land is {Your, My} Land: Evaluating Geopolitical Biases in Language Models

2023-05-24 · Bryan Li, Samar Haider, Chris Callison-Burch

Do the Spratly Islands belong to China, the Philippines, or Vietnam? A pretrained large language model (LLM) may answer differently if asked in the languages of each claimant country: Chinese, Tagalog, or Vietnamese. This contrasts with a multilingual human, who would likely answer consistently. In this paper, we show that LLMs recall certain geographical knowledge inconsistently when queried in different languages -- a phenomenon we term geopolitical bias. As a targeted case study, we consider territorial disputes, an inherently controversial and multilingual task. We introduce BorderLines, a dataset of territorial disputes which covers 251 territories, each associated with a set of multiple-choice questions in the languages of each claimant country (49 languages in total). We also propose a suite of evaluation metrics to precisely quantify bias and consistency in responses across different languages. We then evaluate various multilingual LLMs on our dataset and metrics to probe their internal knowledge and use the proposed metrics to discover numerous inconsistencies in how these models respond in different languages. Finally, we explore several prompt modification strategies, aiming to either amplify or mitigate geopolitical bias, which highlights how brittle LLMs are and how they tailor their responses depending on cues from the interaction context. Our code and data are available at https://github.com/manestay/borderlines

📄 PDF Abstract BibTeX arXiv:2305.14610

Code (1)

manestay/borderlines 공식 구현

Tasks

Language ModellingLarge Language ModelMultiple-choice

Methods 이 논문이 사용한 방법론

BASE 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

Is Your LLM Outdated? Evaluating LLMs at Temporal Generalization

2024-05-14 · Chenghao Zhu, Nuo Chen, Yufei Gao, Yunyi Zhang 외

The rapid advancement of Large Language Models (LLMs) highlights the urgent need for evolving evaluation methodologies that keep pace with improvements in language comprehension and information processing. However, tradi…

Investigating Independence vs. Control: Agenda-Setting in Russian News Coverage on Social Media

2022-06-01 · LREC 2022 6 · Annerose Eichel, Gabriella Lapesa, Sabine Schulte im Walde

Agenda-setting is a widely explored phenomenon in political science: powerful stakeholders (governments or their financial supporters) have control over the media and set their agenda: political and economical powers det…

Strategic Coercion Within Alliances: The Greenland Sovereignty Game as an AI Stress Test

2026-05-11 · Rommin Adl, Peyton Williams arxiv

What happens when the strongest alliance member pressures a weaker member over territory and strategic control? We examine the Greenland sovereignty crisis as a stress test for LLM geopolitics, centered on the 2019-2026 …

A Global Analysis of Cyber Threats to the Energy Sector: "Currents of Conflict" from a Geopolitical Perspective

2025-09-26 · Gustavo Sánchez, Ghada Elbez, Veit Hagenmeyer arxiv

The escalating frequency and sophistication of cyber threats increased the need for their comprehensive understanding. This paper explores the intersection of geopolitical dynamics, cyber threat intelligence analysis, an…

Are we Doomed to an AI Race? Why Self-Interest Could Drive Countries Towards a Moratorium on Superintelligence

2026-05-02 · Edward Roussel, Lode Lauwaert, Torben Swoboda, Grant Ramsey 외 arxiv

This paper uses game theory to argue that, contrary to the prevailing view, a moratorium on Artificial Superintelligence (ASI) can be in a state's self-interest. By formalizing trategic interactions between geopolitical …