You reap what you sow: On the Challenges of Bias Evaluation Under Multilingual Settings
Evaluating bias, fairness, and social impact in monolingual language models is a difficult task. This challenge is further compounded when language modeling occurs in a multilingual context. Considering the implication of evaluation biases for large multilingual language models, we situate the discussion of bias evaluation within a wider context of social scientific research with computational work.We highlight three dimensions of developing multilingual bias evaluation frameworks: (1) increasing transparency through documentation, (2) expanding targets of bias beyond gender, and (3) addressing cultural differences that exist between languages.We further discuss the power dynamics and consequences of training large language models and recommend that researchers remain cognizant of the ramifications of developing such technologies.
Code (0)
등록된 구현이 없습니다.
Tasks
FairnessLanguage ModelingLanguage ModellingSimilar Papers 제목 키워드 기반
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
Video generation models aspire to simulate dynamic environments, and several benchmarks now evaluate memory consistency across frames. However, most assess consistency only while the target remains in view, and the few t…
Video GenerationInteractive Evaluation Requires a Design Science
AI evaluation is undergoing a structural change. Large language models (LLMs) are increasingly deployed as systems that act over time through tools, environments, users, and other agents, while many evaluation practices …
The Benchmark Illusion: Pruned LLMs Can Pass Multiple Choice but Fail to Answer
Compressing large language models reduces memory use and inference cost, but it can also create failures that standard benchmarks miss. A pruned model may still perform well on multiple-choice evaluations, yet fail to an…
Question AnsweringResting-State fingerprints of Acceptance and Reappraisal. The role of Sensorimotor, Executive and Affective networks
Acceptance and Reappraisal are considered adaptive emotion regulation strategies. While previous studies have explored the neural underpinnings of these strategies using task based fMRI and sMRI, a gap exists in the lite…
Functional ConnectivityREAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage
Production deployment of AI coding agents requires fast, reproducible evaluation signals. Existing industrial practices trade off speed and fidelity: online A/B testing takes weeks and risks user experience, shadow deplo…