paper-with-me

Papers

WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models

2025-05-14 · Abdullah Mushtaq, Imran Taj, Rafay Naeem, Ibrahim Ghaznavi, Junaid Qadir

Large Language Models (LLMs) are predominantly trained and aligned in ways that reinforce Western-centric epistemologies and socio-cultural norms, leading to cultural homogenization and limiting their ability to reflect global civilizational plurality. Existing benchmarking frameworks fail to adequately capture this bias, as they rely on rigid, closed-form assessments that overlook the complexity of cultural inclusivity. To address this, we introduce WorldView-Bench, a benchmark designed to evaluate Global Cultural Inclusivity (GCI) in LLMs by analyzing their ability to accommodate diverse worldviews. Our approach is grounded in the Multiplex Worldview proposed by Senturk et al., which distinguishes between Uniplex models, reinforcing cultural homogenization, and Multiplex models, which integrate diverse perspectives. WorldView-Bench measures Cultural Polarization, the exclusion of alternative perspectives, through free-form generative evaluation rather than conventional categorical benchmarks. We implement applied multiplexity through two intervention strategies: (1) Contextually-Implemented Multiplex LLMs, where system prompts embed multiplexity principles, and (2) Multi-Agent System (MAS)-Implemented Multiplex LLMs, where multiple LLM agents representing distinct cultural perspectives collaboratively generate responses. Our results demonstrate a significant increase in Perspectives Distribution Score (PDS) entropy from 13% at baseline to 94% with MAS-Implemented Multiplex LLMs, alongside a shift toward positive sentiment (67.7%) and enhanced cultural balance. These findings highlight the potential of multiplex-aware AI evaluation in mitigating cultural bias in LLMs, paving the way for more inclusive and ethically aligned AI systems.

📄 PDF Abstract BibTeX arXiv:2505.09595

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

CCD-Bench: Probing Cultural Conflict in Large Language Model Decision-Making

2025-10-03 · Hasibur Rahman, Hanan Salam arxiv

Although large language models (LLMs) are increasingly implicated in interpersonal and societal decision-making, their ability to navigate explicit conflicts between legitimately different cultural value systems remains …

Value predictionDecision MakingBias Detection

NRITYAM: Language Models Meet Art and Heritage of Dance

2026-06-18 · Punit Kumar Singh, Niladri Ghosh, Advait Joshiınst, Shailee Choudhary 외 arxiv

Language models have become essential tools in shaping modern workflows. However, their global effectiveness hinges on a nuanced understanding of local socio-cultural contexts. To address this gap, we present NRITYAM, a …

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages

2024-11-25 · CVPR 2025 1 · Ashmal Vayani, Dinura Dissanayake, Hasindri Watawana, Noor Ahsan 외

Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cultural contexts, respect local sensitivit…

AllLong Question AnswerMultiple-choiceShort Question Answers+1

Culture in Action: Evaluating Text-to-Image Models through Social Activities

2025-11-07 · Sina Malakouti, Boqing Gong, Adriana Kovashka arxiv

Text-to-image (T2I) diffusion models achieve impressive photorealism by training on large-scale web data, but models inherit cultural biases and fail to depict underrepresented regions faithfully. Existing cultural bench…

The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models

2026-04-22 · Yilun Liu, Chunguang Zhao, Mengyao Piao, Lingqi Miao 외 arxiv

Evaluating the multilingual and multicultural capabilities of Large Language Models (LLMs) is essential for their global utility. However, current benchmarks face three critical limitations: (1) fragmented evaluation dim…

Machine Translation