paper-with-me

Papers

Whatever Remains Must Be True: Filtering Drives Reasoning in LLMs, Shaping Diversity

2025-12-05 · Germán Kruszewski, Pierre Erbacher, Jos Rozen, Marc Dymetman arxiv

Reinforcement Learning (RL) has become the de facto standard for tuning LLMs to solve tasks involving reasoning. However, growing evidence shows that models trained in such way often suffer from a significant loss in diversity. We argue that this arises because RL implicitly optimizes the "mode-seeking" or "zero-forcing" Reverse KL to a target distribution causing the model to concentrate mass on certain high-probability regions of the target while neglecting others. In this work, we instead begin from an explicit target distribution, obtained by filtering out incorrect answers while preserving the relative probabilities of correct ones. Starting from a pre-trained LLM, we approximate this target distribution using the $α$-divergence family, which unifies prior approaches and enables direct control of the precision-diversity trade-off by interpolating between mode-seeking and mass-covering divergences. On a Lean theorem-proving benchmark, our method achieves state-of-the-art performance along the coverage-precision Pareto frontier, outperforming all prior methods on the coverage axis.

📄 PDF Abstract BibTeX arXiv:2512.05962

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Illustrating a neural model of logic computations: The case of Sherlock Holmes' old maxim

2012-10-28 · Eduardo Mizraji

Natural languages can express some logical propositions that humans are able to understand. We illustrate this fact with a famous text that Conan Doyle attributed to Holmes: 'It is an old maxim of mine that when you have…

Eliminating The Impossible, Whatever Remains Must Be True

2022-06-20 · Jinqiang Yu, Alexey Ignatiev, Peter J. Stuckey, Nina Narodytska 외

The rise of AI methods to make predictions and decisions has led to a pressing need for more explainable artificial intelligence (XAI) methods. One common approach for XAI is to produce a post-hoc explanation, explaining…

Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)Prediction

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

2026-07-30 · Jin Cao, Zian Meng, Kaipeng Zhang arxiv

We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existing interfaces either encode an action loosely, leaving how it unfold…

DriveSuprim: Towards Precise Trajectory Selection for End-to-End Planning

2025-06-07 · Wenhao Yao, Zhenxin Li, Shiyi Lan, Zi Wang 외

In complex driving environments, autonomous vehicles must navigate safely. Relying on a single predicted path, as in regression-based approaches, usually does not explicitly assess the safety of the predicted trajectory.…

Autonomous VehiclesCollision AvoidanceNavigateNavSim

Classical Limits of Spectral Filtering in Quantum Generative Models

2026-08-14 · Marco Roth arxiv

Spectral filtering has been proposed as a route to regularization in quantum generative models: the quantum Fourier transform exposes the amplitude spectrum of a quantum circuit Born machine, and a diagonal filter suppre…