paper-with-me

홈 › Papers

Evaluating Cultural and Social Awareness of LLM Web Agents

2024-10-30 · Haoyi Qiu, Alexander R. Fabbri, Divyansh Agarwal, Kung-Hsiang Huang, Sarah Tan, Nanyun Peng, Chien-Sheng Wu

As large language models (LLMs) expand into performing as agents for real-world applications beyond traditional NLP tasks, evaluating their robustness becomes increasingly important. However, existing benchmarks often overlook critical dimensions like cultural and social awareness. To address these, we introduce CASA, a benchmark designed to assess LLM agents' sensitivity to cultural and social norms across two web-based tasks: online shopping and social discussion forums. Our approach evaluates LLM agents' ability to detect and appropriately respond to norm-violating user queries and observations. Furthermore, we propose a comprehensive evaluation framework that measures awareness coverage, helpfulness in managing user queries, and the violation rate when facing misleading web content. Experiments show that current LLMs perform significantly better in non-agent than in web-based agent environments, with agents achieving less than 10% awareness coverage and over 40% violation rates. To improve performance, we explore two methods: prompting and fine-tuning, and find that combining both methods can offer complementary advantages -- fine-tuning on culture-specific datasets significantly enhances the agents' ability to generalize across different regions, while prompting boosts the agents' ability to navigate complex tasks. These findings highlight the importance of constantly benchmarking LLM agents' cultural and social awareness during the development cycle.

📄 PDF Abstract BibTeX arXiv:2410.23252

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingNavigate

Similar Papers 제목 키워드 기반

Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment

2025-09-26 · Hongbin Zhang, Kehai Chen, Xuefeng Bai, Yang Xiang 외 arxiv

Reward models (RMs) are crucial for aligning large language models (LLMs) with diverse cultures. Consequently, evaluating their cultural awareness is essential for further advancing global alignment of LLMs. However, exi…

Reinforcement Learning

DaKultur: Evaluating the Cultural Awareness of Language Models for Danish with Native Speakers

2025-04-03 · Max Müller-Eberstein, Mike Zhang, Elisa Bassignana, Peter Brunsgaard Trolle 외

Large Language Models (LLMs) have seen widespread societal adoption. However, while they are able to interact with users in languages beyond English, they have been shown to lack cultural awareness, providing anglocentri…

Survey of Cultural Awareness in Language Models: Text and Beyond

2024-10-30 · Siddhesh Pawar, Junyeong Park, Jiho Jin, Arnav Arora 외

Large-scale deployment of large language models (LLMs) in various applications, such as chatbots and virtual assistants, requires LLMs to be culturally sensitive to the user to ensure inclusivity. Culture has been widely…

Benchmarking

ExCAM: Explainable Cultural Awareness Metrics

2026-05-28 · Christoph Leiter, Haiyue Song, Hour Kaing, Jin Tei 외 arxiv

Evaluating the cultural awareness of large language models is crucial to ensure the fairness of generated text and the generalizability of applications across the world. Recent benchmarks explore cultural goods like food…

Question AnsweringText Generation

WorldValuesBench: A Large-Scale Benchmark Dataset for Multi-Cultural Value Awareness of Language Models

2024-04-25 · Wenlong Zhao, Debanjan Mondal, Niket Tandon, Danica Dillion 외

The awareness of multi-cultural human values is critical to the ability of language models (LMs) to generate safe and personalized responses. However, this awareness of LMs has been insufficiently studied, since the comp…

Value prediction