paper-with-me

Papers

You reap what you sow: On the Challenges of Bias Evaluation Under Multilingual Settings

2022-05-01 · BigScience (ACL) 2022 5 · Zeerak Talat, Aurélie Névéol, Stella Biderman, Miruna Clinciu, Manan Dey, Shayne Longpre, Sasha Luccioni, Maraim Masoud, Margaret Mitchell, Dragomir Radev, Shanya Sharma, Arjun Subramonian, Jaesung Tae, Samson Tan, Deepak Tunuguntla, Oskar van der Wal

Evaluating bias, fairness, and social impact in monolingual language models is a difficult task. This challenge is further compounded when language modeling occurs in a multilingual context. Considering the implication of evaluation biases for large multilingual language models, we situate the discussion of bias evaluation within a wider context of social scientific research with computational work.We highlight three dimensions of developing multilingual bias evaluation frameworks: (1) increasing transparency through documentation, (2) expanding targets of bias beyond gender, and (3) addressing cultural differences that exist between languages.We further discuss the power dynamics and consequences of training large language models and recommend that researchers remain cognizant of the ramifications of developing such technologies.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

FairnessLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments

2026-06-25 · Haoyu Chen, Kaichen Zhou, Hang Hua, Kaile Zhang 외 arxiv

Video generation models aspire to simulate dynamic environments, and several benchmarks now evaluate memory consistency across frames. However, most assess consistency only while the target remains in view, and the few t…

Video Generation

Interactive Evaluation Requires a Design Science

2026-05-18 · Keyang Xuan, Peiyang Song, Pan Lu, Pengrui Han 외 arxiv

AI evaluation is undergoing a structural change. Large language models (LLMs) are increasingly deployed as systems that act over time through tools, environments, users, and other agents, while many evaluation practices …

The Benchmark Illusion: Pruned LLMs Can Pass Multiple Choice but Fail to Answer

2026-06-16 · Rui Wen, Lu Sun, Jiayang Liu, Zesheng Xu 외 arxiv

Compressing large language models reduces memory use and inference cost, but it can also create failures that standard benchmarks miss. A pruned model may still perform well on multiple-choice evaluations, yet fail to an…

Question Answering

Resting-State fingerprints of Acceptance and Reappraisal. The role of Sensorimotor, Executive and Affective networks

2024-01-29 · Parisa Ahmadi Ghomroudi, Roma Siugzdaite, Irene Messina, Alessandro Grecucci

Acceptance and Reappraisal are considered adaptive emotion regulation strategies. While previous studies have explored the neural underpinnings of these strategies using task based fMRI and sMRI, a gap exists in the lite…

Functional Connectivity

REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage

2026-04-02 · Smriti Jha, Matteo Paltenghi, Chandra Maddila, Vijayaraghavan Murali 외 arxiv

Production deployment of AI coding agents requires fast, reproducible evaluation signals. Existing industrial practices trade off speed and fidelity: online A/B testing takes weeks and risks user experience, shadow deplo…