paper-with-me

홈 › Papers

Are LLMs Ready for Conflict Monitoring? Empirical Evidence from West Africa

2026-05-05 · Hoffmann Muki, Olukunle Owolabi arxiv

As LLMs enter conflict monitoring, understanding systematic distortions in their outputs is critical for humanitarian accountability. We evaluate four vanilla open-weight models Gemma 3 4B, Llama 3.2 3B, Mistral 7B, and OLMo 2 7B and two domain-adapted models, AfroConfliBERT and AfroConfliLLAMA, on Nigeria and Cameroon conflict-event classification against ACLED, a gold-standard dataset with multi-stage verification. We find a bifurcated divergence in normative directionality. Open-weight models exhibit statistically significant False Illegitimation bias: Gemma misclassifies to 18.29% of legitimate battles as civilian-targeted violence while making zero False Legitimation errors. By contrast, AfroConfliBERT and AfroConfliLLAMA achieve near-directional neutrality, with Legitimization Bias differences indistinguishable from zero. Yet domain adaptation does not eliminate actor-based selection bias. Both adapted models show statistically significant actor bias comparable to vanilla LLMs; in Nigeria, state actors are legitimized 36.5% more often than non-state actors in identical tactical contexts. Open-weight outputs are also fragile to geography-specific lexical framing: delegitimizing phrases produce flip rates up to 66.7% in Cameroon and 34.2% in Nigeria, while perturbations salient in one context may not matter in another. Error trace profiling shows models mask normative bias through unfaithful rationale confabulations. In contrast, AfroConfliBERT and AfroConfliLLAMA are largely robust, with near-zero flip rates across perturbation categories. Overall, current models are not ready for unsupervised deployment in conflict monitoring. We call for fairness-aware fine-tuning to reduce actor-based selection bias, mandatory adversarial robustness evaluation against lexical manipulation, and context-specific human-in-the-loop oversight calibrated to regional difficulty.

📄 PDF Abstract BibTeX arXiv:2605.04177

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial RobustnessDomain Adaptation

Similar Papers 제목 키워드 기반

Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts

2026-09-05 · Yen-Ting Piao, Shu-Yun Chen, Chin-Hui Chu, Chun-Wei Chen 외 hf

Omni-modal large language models (OLLMs) jointly process vision, audio, and text, yet their modality bias under cross-modal conflict remains underexplored. Existing benchmarks conflate two distinct forms of evidence with…

ECon: On the Detection and Resolution of Evidence Conflicts

2024-10-05 · Cheng Jiayang, Chunkit Chan, Qianqian Zhuang, Lin Qiu 외

The rise of large language models (LLMs) has significantly influenced the quality of information in decision-making systems, leading to the prevalence of AI-generated content and challenges in detecting misinformation an…

Decision MakingMisinformationNatural Language Inference

Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge Conflicts

2023-05-22 · Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou 외

By providing external information to large language models (LLMs), tool augmentation (including retrieval augmentation) has emerged as a promising solution for addressing the limitations of LLMs' static parametric memory…

Retrieval

Exploring Knowledge Conflicts for Faithful LLM Reasoning: Benchmark and Method

2026-04-13 · Tianzhe Zhao, Jiaoyan Chen, Shuxiu Zhang, Haiping Zhu 외 arxiv

Large language models (LLMs) have achieved remarkable success across a wide range of applications especially when augmented by external knowledge through retrieval-augmented generation (RAG). Despite their widespread ado…

Knowledge Graphs

Industrial AI Robustness Card for Time Series Models

2025-12-05 · Alexander Windmann, Benedikt Stratmann, Mariya Lyashenko, Oliver Niggemann arxiv

Industrial AI practitioners face vague robustness requirements in emerging regulations and standards but lack concrete, implementation-ready protocols. This paper introduces the Industrial AI Robustness Card for Time Ser…