The Subjectivity of Monoculture
Machine learning models -- including large language models (LLMs) -- are often said to exhibit monoculture, where outputs agree strikingly often. But what does it actually mean for models to agree too much? We argue that this question is inherently subjective, relying on two key decisions. First, the analyst must specify a baseline null model for what "independence" should look like. This choice is inherently subjective, and as we show, different null models result in dramatically different inferences about excess agreement. Second, we show that inferences depend on the population of models and items under consideration. Models that seem highly correlated in one context may appear independent when evaluated on a different set of questions, or against a different set of peers. Experiments on two large-scale benchmarks validate our theoretical findings. For example, we find drastically different inferences when using a null model with item difficulty compared to previous works that do not. Together, our results reframe monoculture evaluation not as an absolute property of model behavior, but as a context-dependent inference problem.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Strategic Algorithmic Monoculture: Experimental Evidence from Coordination Games
AI agents increasingly operate in multi-agent environments where outcomes depend on coordination. We distinguish primary algorithmic monoculture -- baseline action similarity -- from strategic algorithmic monoculture, wh…
Generative Monoculture in Large Language Models
We introduce {\em generative monoculture}, a behavior observed in large language models (LLMs) characterized by a significant narrowing of model output diversity relative to available training data for a given task: for …
Code GenerationDiversityAlgorithmic Monoculture and Social Welfare
As algorithms are increasingly applied to screen applicants for high-stakes decisions in employment, lending, and other domains, concerns have been raised about the effects of algorithmic monoculture, in which many decis…
Decision MakingFrom Protoscience to Epistemic Monoculture: How Benchmarking Set the Stage for the Deep Learning Revolution
Over the past decade, AI research has focused heavily on building ever-larger deep learning models. This approach has simultaneously unlocked incredible achievements in science and technology, and hindered AI from overco…
BenchmarkingNot even wrong: Reply to Wagg et al
We demonstrate that the issues described in the Wagg et al. (2019) Comment on our paper (Pillai and Gouhier, 2019) are all due to misunderstandings about the implications of pairwise effects, the nature of the null basel…
Diversity