paper-with-me

Papers

Evaluating Stability of Unreflective Alignment

2024-08-27 · James Lucassen, Mark Henry, Philippa Wright, Owen Yeung

Many theoretical obstacles to AI alignment are consequences of reflective stability - the problem of designing alignment mechanisms that the AI would not disable if given the option. However, problems stemming from reflective stability are not obviously present in current LLMs, leading to disagreement over whether they will need to be solved to enable safe delegation of cognitive labor. In this paper, we propose Counterfactual Priority Change (CPC) destabilization as a mechanism by which reflective stability problems may arise in future LLMs. We describe two risk factors for CPC-destabilization: 1) CPC-based stepping back and 2) preference instability. We develop preliminary evaluations for each of these risk factors, and apply them to frontier LLMs. Our findings indicate that in current LLMs, increased scale and capability are associated with increases in both CPC-based stepping back and preference instability, suggesting that CPC-destabilization may cause reflective stability problems in future LLMs.

📄 PDF Abstract BibTeX arXiv:2408.15116

Code (0)

등록된 구현이 없습니다.

Tasks

counterfactual

Similar Papers 제목 키워드 기반

Lazy Data Practices Harm Fairness Research

2024-04-26 · Jan Simson, Alessandro Fabris, Christoph Kern

Data practices shape research and practice on fairness in machine learning (fair ML). Critical data studies offer important reflections and critiques for the responsible advancement of the field by highlighting shortcomi…

Fairness

Fairness Incentives in Response to Unfair Dynamic Pricing

2024-04-22 · Jesse Thibodeau, Hadi Nekoei, Afaf Taïk, Janarthanan Rajendran 외

The use of dynamic pricing by profit-maximizing firms gives rise to demand fairness concerns, measured by discrepancies in consumer groups' demand responses to a given pricing strategy. Notably, dynamic pricing may resul…

FairnessReinforcement Learning (RL)

Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces

2026-05-17 · Seth Karten, Cameron Crow, Chi Jin arxiv

The deployment of Large Language Models (LLMs) as autonomous economic agents introduces systemic risks that extend beyond individual capability failures. As agents transition to directly interacting with marketplaces, th…

Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability

2026-02-19 · Nusrat Jahan Lia, Shubhashis Roy Dipta arxiv

Recent advances in multilingual representation learning aim to bridge the performance gap between high- and low-resource languages, yet their ability to preserve affective meaning across languages remains underexplored, …

Representation Learning

EMPA: Evaluating Persona-Aligned Empathy as a Process

2026-02-28 · Shiya Zhang, Yuhan Zhan, Ruixi Su, Ruihan Sun 외 arxiv

Evaluating persona-aligned empathy in LLM-based dialogue agents remains challenging. User states are latent, feedback is sparse and difficult to verify in situ, and seemingly supportive turns can still accumulate into tr…