paper-with-me

홈 › Papers

Iterative Finetuning is Mostly Idempotent

2026-05-01 · Zephaniah Roe, Jack Sanderson, Dang Nguyen, Julian Huang, Todd Nief, Aryan Shrivastava, Chenhao Tan, Ari Holtzman arxiv

If a model has some behavioral tendency, such as sycophancy or misalignment, and it is trained on its own outputs, will the tendency be amplified in the next generation of models? We study this question by training a series of models where each model is finetuned on data generated by its predecessor, and the initial model is seeded with some persona or belief. We test three settings: supervised finetuning (SFT) on instruct models, synthetic document finetuning (SDF) on base models, and direct preference optimization (DPO). In the SFT and SDF settings, traits mostly decay or remain constant so that further finetuning cycles do nothing. In rare cases when amplification occurs, it generally comes at the cost of coherence. In the DPO setting, trait amplification can reliably occur when a model is continually trained with a preference for its own outputs, but vanishes when models are reinitialized at each cycle. Overall, our results suggest that amplification most likely comes from continual post-training, and limiting this stage may be an effective defense. For non-RL finetuning, trait amplification is rare and very sensitive to data quantity, making it significantly less likely to occur accidentally. Finally, the amplification-coherence tradeoff serves as a natural deterrent against trait amplification.

📄 PDF Abstract BibTeX arXiv:2605.01130

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Score-based Idempotent Distillation of Diffusion Models

2025-09-25 · Shehtab Zaman, Chengyan Liu, Kenneth Chiu arxiv

Idempotent generative networks (IGNs) are a new line of generative models based on idempotent mapping to a target manifold. IGNs support both single-and multi-step generation, allowing for a flexible trade-off between co…

Idempotent Learned Image Compression with Right-Inverse

2023-09-21 · NeurIPS 2023 11

We consider the problem of idempotent learned image compression (LIC). The idempotence of codec refers to the stability of codec to re-compression. To achieve idempotence, previous codecs adopt invertible transforms such…

Orthogonal and Idempotent Transformations for Learning Deep Neural Networks

2017-07-19 · Jingdong Wang, Yajie Xing, Kexin Zhang, Cha Zhang

Identity transformations, used as skip-connections in residual networks, directly connect convolutional layers close to the input and those close to the output in deep neural networks, improving information flow and thus…

A* shortest string decoding for non-idempotent semirings

2022-04-14 · Kyle Gorman, Cyril Allauzen

The single shortest path algorithm is undefined for weighted finite-state automata over non-idempotent semirings because such semirings do not guarantee the existence of a shortest path. However, in non-idempotent semiri…

IDEM Enough? Evolving Highly Nonlinear Idempotent Boolean Functions

2026-01-31 · Claude Carlet, Marko Ðurasevic, Domagoj Jakobovic, Luca Mariot 외 arxiv

Idempotent Boolean functions form a highly structured subclass of Boolean functions that is closely related to rotation symmetry under a normal-basis representation and to invariance under a fixed linear map in a polynom…