paper-with-me

홈 › Papers

Peer-Predictive Self-Training for Language Model Reasoning

2026-04-14 · Shi Feng, Hanlin Zhang, Fan Nie, Sham Kakade, Yiling Chen arxiv

Mechanisms for continued self-improvement of language models without external supervision remain an open challenge. We propose Peer-Predictive Self-Training (PST), a label-free fine-tuning framework in which multiple language models improve collaboratively by using a cross-model aggregate response as an internal training signal. Given a prompt, models generate responses sequentially; the final aggregated answer, which is often more reliable than individual responses in practice, serves as an internal reference for learning. We measure how informative each intermediate response is about the aggregate using pointwise mutual information (PMI), and use this signal to scale self-training updates: responses already aligned with the aggregate receive smaller updates, while less informative or misaligned responses receive larger ones. On mathematical reasoning benchmarks, including SimulEq, MATH-500-Numeric, and MultiArith, PST improves exact-match accuracy by 2.2--4.3 percentage points across Gemma-2-2B, LLaMA-3.2-1B, and Qwen2.5-1.5B, and reduces the average generator--verifier gap (GV-Gap) by 26--40%, while requiring no external supervision, no teacher--student hierarchy, and only cross-model interactions. These results suggest that peer-predictive feedback from cross-model generations can provide an effective mechanism for self-supervised language-model improvement.

📄 PDF Abstract BibTeX arXiv:2604.13356

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical Reasoning

Similar Papers 제목 키워드 기반

Towards Reasoning in Large Language Models via Multi-Agent Peer Review Collaboration

2023-11-14 · Zhenran Xu, Senbao Shi, Baotian Hu, Jindi Yu 외

Large Language Models (LLMs) have shown remarkable capabilities in general natural language processing tasks but often fall short in complex reasoning tasks. Recent studies have explored human-like problem-solving strate…

DiversityMath

Learning from Peers in Reasoning Models

2025-05-12 · Tongxu Luo, Wenyu Du, Jiaxi Bi, Stephen Chung 외

Large Reasoning Models (LRMs) have the ability to self-correct even when they make mistakes in their reasoning paths. However, our study reveals that when the reasoning process starts with a short but poor beginning, it …

Math

When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning

2025-10-08 · Hyeong Kyu Choi, Xiaojin Zhu, Sharon Li arxiv

Multi-agent debate (MAD) aims to improve large language model (LLM) reasoning by letting multiple agents exchange answers and then aggregate their opinions. Yet recent studies reveal that agents are not neutral: they are…

Localizing Credit at the Divergence: Path-Conditioned Self-Distillation for LLM Reasoning

2026-06-14 · Yu Li, Shu Hong, Tian Lan arxiv

Reinforcement learning from verifiable rewards assigns a single scalar to each rollout, leaving token-level credit assignment underspecified in long reasoning traces. On-policy self-distillation addresses this by letting…

Reinforcement Learning

PEER: A Collaborative Language Model

2022-08-24 · Timo Schick, Jane Dwivedi-Yu, Zhengbao Jiang, Fabio Petroni 외

Textual content is often the output of a collaborative writing process: We start with an initial draft, ask for suggestions, and repeatedly make changes. Agnostic of this process, today's language models are trained to g…

DiversityLanguage ModelingLanguage Modellingmodel