paper-with-me

Papers

Weird Generalization is Weirdly Brittle

2026-04-11 · Miriam Wanner, Hannah Collison, William Jurayj, Benjamin Van Durme, Mark Dredze, William Walden arxiv

Weird generalization is a phenomenon in which models fine-tuned on data from a narrow domain (e.g. insecure code) develop surprising traits that manifest even outside that domain (e.g. broad misalignment)-a phenomenon that prior work has highlighted as a critical safety concern. Here, we present an extended replication study of key weird generalization results across an expanded suite of models and datasets. We confirm that surprising (and dangerous) traits can emerge under certain circumstances, but we find that weird generalization is exceptionally brittle: it emerges only for specific models on specific datasets, and it vanishes under simple training-time, prompt-based interventions. We find that the most effective interventions provide prompt context that makes the generalized behavior the expected behavior. However, we show that even very generic interventions that do not anticipate specific generalized traits can still be effective in mitigating weird generalization's effects. Our findings thus help clarify the nature of the safety threat that weird generalization poses and point toward an easily implemented set of solutions.

📄 PDF Abstract BibTeX arXiv:2604.10022

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Discipline and Label: A WEIRD Genealogy and Social Theory of Data Annotation

2024-02-09 · Andrew Smart, Ding Wang, Ellis Monk, Mark Díaz 외

Data annotation remains the sine qua non of machine learning and AI. Recent empirical work on data annotation has begun to highlight the importance of rater diversity for fairness, model performance, and new lines of res…

Fairness

AutoWeird: Weird Translational Scoring Function Identified by Random Search

2022-07-24 · Hansi Yang, Yongqi Zhang, Quanming Yao

Scoring function (SF) measures the plausibility of triplets in knowledge graphs. Different scoring functions can lead to huge differences in link prediction performances on different knowledge graphs. In this report, we …

AttributeKnowledge GraphsLink PredictionTriplet

On the Threat Model of Weird Generalization and Emergent Misalignment

2026-08-24 · Miriam Wanner, Mark Dredze, William Walden arxiv

Narrow fine-tuning on small, domain-specific datasets can produce broad and surprising changes in model behavior-a phenomenon called weird generalization (WG). Yet, it remains unclear what features of the fine-tuning dat…

Should LLMs be WEIRD? Exploring WEIRDness and Human Rights in Large Language Models

2025-08-22 · Ke Zhou, Marios Constantinides, Daniele Quercia arxiv

Large language models (LLMs) are often trained on data that reflect WEIRD values: Western, Educated, Industrialized, Rich, and Democratic. This raises concerns about cultural bias and fairness. Using responses to the Wor…

WEIRD ICWSM: How Western, Educated, Industrialized, Rich, and Democratic is Social Computing Research?

2024-06-04 · Ali Akbar Septiandri, Marios Constantinides, Daniele Quercia

Much of the research in social computing analyzes data from social media platforms, which may inherently carry biases. An overlooked source of such bias is the over-representation of WEIRD (Western, Educated, Industriali…

Diversity