paper-with-me

홈 › Papers

IRIS: Interpolative Rényi Iterative Self-play for Large Language Model Fine-Tuning

2026-04-22 · Wenjie Liao, Like Wu, Liangjie Zhao, Shihui Xu, Shigeru Fujimura arxiv

Self-play fine-tuning enables large language models to improve beyond supervised fine-tuning without additional human annotations by contrasting annotated responses with self-generated ones. Many existing methods rely on a fixed divergence regime. SPIN is closely related to a KL-based regime, SPACE to a Jensen-Shannon-style objective via noise contrastive estimation, and SPIF to $χ^2$-regularized self-play. Since these divergences exhibit different strengths depending on the distributional gap between model and target, no single choice appears to provide favorable learning dynamics across training stages. We propose IRIS (Interpolative Rényi Iterative Self-play), a Rényi-based self-play fine-tuning framework with a continuously adjustable objective. IRIS decomposes into two independent tilted risk terms over annotated and synthetic data, with exponential importance weights controlled by the order parameter $α$. We show that several self-play objectives can be interpreted as limiting or representative regimes at particular values of $α$, providing a unified theoretical perspective on these methods. An adaptive order schedule further adjusts $α$ to the distributional gap, shifting from sharper importance weighting early in training to smoother refinement near convergence. Theoretically, we establish the fixed-point property of IRIS and analyze how $α$ controls gradient concentration. Experiments on Zephyr-7B and Qwen2.5-3B across ten benchmarks show that IRIS improves upon baselines, reaching 44.57\% average score with gains across iterations. In our setting, IRIS with only 26$k$ annotated samples surpasses standard supervised fine-tuning trained on the full 200$k$ dataset.

📄 PDF Abstract BibTeX arXiv:2604.20933

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation

2024-05-21 · Govind Ramesh, Yao Dou, Wei Xu

Research on jailbreaking has been valuable for testing and understanding the safety and security issues of large language models (LLMs). In this paper, we introduce Iterative Refinement Induced Self-Jailbreak (IRIS), a n…

Interpolative Decoding: Exploring the Spectrum of Personality Traits in LLMs

2025-12-23 · Eric Yeh, John Cadigan, Ran Chen, Dick Crouch 외 arxiv

Recent research has explored using very large language models (LLMs) as proxies for humans in tasks such as simulation, surveys, and studies. While LLMs do not possess a human psychology, they often can emulate human beh…

Decision Making

Graph-Informed Adversarial Modeling: Infimal Subadditivity of Interpolative Divergences

2026-03-20 · Panagiota Birmpa, Eric Joseph Hall arxiv

We study adversarial learning when the target distribution factorizes according to a known Bayesian network. For interpolative divergences, including $(f,Γ)$-divergences, we prove a new infimal subadditivity principle sh…

Post-Mortem Iris Recognition Resistant to Biological Eye Decay Processes

2019-12-05 · Mateusz Trokielewicz, Adam Czajka, Piotr Maciejewicz

This paper proposes an end-to-end iris recognition method designed specifically for post-mortem samples, and thus serving as a perfect application for iris biometrics in forensics. To our knowledge, it is the first metho…

Iris Recognition

Spoofing PRNU Patterns of Iris Sensors while Preserving Iris Recognition

2018-08-31 · Sudipta Banerjee, Vahid Mirjalili, Arun Ross

The principle of Photo Response Non-Uniformity (PRNU) is used to link an image with its source, i.e., the sensor that produced it. In this work, we investigate if it is possible to modify an iris image acquired using one…

Iris Recognition