paper-with-me

홈 › Papers

Bad Seeds: Evaluating Lexical Methods for Bias Measurement

2021-08-01 · ACL 2021 5 · Maria Antoniak, David Mimno

A common factor in bias measurement methods is the use of hand-curated seed lexicons, but there remains little guidance for their selection. We gather seeds used in prior work, documenting their common sources and rationales, and in case studies of three English-language corpora, we enumerate the different types of social biases and linguistic features that, once encoded in the seeds, can affect subsequent bias measurements. Seeds developed in one context are often re-used in other contexts, but documentation and evaluation remain necessary precursors to relying on seeds for sensitive measurements.

📄 PDF Abstract BibTeX

Code (1)

maria-antoniak/bad-seeds 공식 구현

Similar Papers 제목 키워드 기반

[Re] Badder Seeds: Reproducing the Evaluation of Lexical Methods for Bias Measurement

2022-06-03 · Jille van der Togt, Lea Tiyavorabun, Matteo Rosati, Giulio Starace

Combating bias in NLP requires bias measurement. Bias measurement is almost always achieved by using lexicons of seed terms, i.e. sets of words specifying stereotypes or dimensions of interest. This reproducibility study…

Assessing the Reliability of Word Embedding Gender Bias Measures

2021-09-10 · EMNLP 2021 11 · Yupei Du, Qixiang Fang, Dong Nguyen

Various measures have been proposed to quantify human-like social biases in word embeddings. However, bias scores based on these measures can suffer from measurement error. One indication of measurement quality is reliab…

Word Embeddings

SURGELLM: Rethinking Multi-Task Evaluation through Task-Aware Feature Gating with Class-Balanced Normalization

2026-06-23 · Noor Islam S. Mohammad, Ulug Bayazit arxiv

Fine-tuned encoders deployed across heterogeneous NLP tasks face three compounding problems: mismatched inductive biases, class-imbalance corruption of feature statistics, and no mechanism to condition attention on exter…

Robust Bias Evaluation with FilBBQ: A Filipino Bias Benchmark for Question-Answering Language Models

2026-02-16 · Lance Calvin Lim Gamboa, Yue Feng, Mark Lee arxiv

With natural language generation becoming a popular use case for language models, the Bias Benchmark for Question-Answering (BBQ) has grown to be an important benchmark format for evaluating stereotypical associations ex…

Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditions

2026-07-01 · Zewen Liu arxiv

The bias-reliability tradeoff conjectures that LLM evaluation systems are constrained in (gamma, H, CV) space, where evaluator coupling (gamma), strategy diversity (H), and small-sample measurement reliability (CV(N)) ca…