paper-with-me

홈 › Papers

CORE: Collaborative Reasoning via Cross Teaching

2026-01-29 · Kshitij Mishra, Mirat Aubakirov, Martin Takac, Nils Lukas, Salem Lahlou arxiv

Large language models exhibit complementary reasoning errors: on the same instance, one model may succeed with a particular decomposition while another fails. We propose Collaborative Reasoning (CORE), a training-time collaboration framework that converts peer success into a learning signal via a cross-teaching protocol. Each problem is solved in two stages: a cold round of independent sampling, followed by a contexted rescue round in which models that failed receive hint extracted from a successful peer. CORE optimizes a combined reward that balances (i) correctness, (ii) a lightweight DPP-inspired diversity term to reduce error overlap, and (iii) an explicit rescue bonus for successful recovery. We evaluate CORE across four standard reasoning datasets GSM8K, MATH, AIME, and GPQA. With only 1,000 training examples, a pair of small open source models (3B+4B) reaches Pass@2 of 99.54% on GSM8K and 92.08% on MATH, compared to 82.50% and 74.82% for single-model training. On harder datasets, the 3B+4B pair reaches Pass@2 of 77.34% on GPQA (trained on 348 examples) and 79.65% on AIME (trained on 792 examples), using a training-time budget of at most 1536 context tokens and 3072 generated tokens. Overall, these results show that training-time collaboration can reliably convert model complementarity into large gains without scaling model size.

📄 PDF Abstract BibTeX arXiv:2601.21600

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Collaborative and Privacy-Preserving Machine Teaching via Consensus Optimization

2019-05-07 · Yufei Han, Yuzhe ma, Christopher Gates, Kevin Roundy 외

In this work, we define a collaborative and privacy-preserving machine teaching paradigm with multiple distributed teachers. We focus on consensus super teaching. It aims at organizing distributed teachers to jointly sel…

Privacy Preserving

Learning from Diverse Reasoning Paths with Routing and Collaboration

2025-08-23 · Zhenyu Lei, Zhen Tan, Song Wang, Yaochen Zhu 외 arxiv

Advances in large language models (LLMs) significantly enhance reasoning capabilities but their deployment is restricted in resource-constrained scenarios. Knowledge distillation addresses this by transferring knowledge …

Knowledge Distillation

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code

2025-07-10 · Keqin Bao, Nuo Chen, Xiaoyuan Li, Binyuan Hui 외 arxiv

Enhancing reasoning capabilities remains a central focus in the LLM reasearch community. A promising direction involves requiring models to simulate code execution step-by-step to derive outputs for given inputs. However…

Reinforcement LearningLogical Reasoning

AI-Driven Analytics of Team-Teaching Talk: Acoustic Patterns across Experience, Cohorts and the Learning Design

2026-04-19 · Yuchen Liu, Roberto Martinez-Maldonado, Riordan Alfredo, Paola Mejia-Domenzain 외 arxiv

As classroom cohorts expand, team teaching is increasingly used to integrate the expertise and pedagogical perspectives of multiple teachers. Yet, there is limited empirical understanding of how team teaching unfolds in …

Joint Video Summarization and Moment Localization by Cross-Task Sample Transfer

2022-01-01 · CVPR 2022 1 · Hao Jiang, Yadong Mu

Video summarization has recently engaged increasing attention in computer vision communities. However, the scarcity of annotated data has been a key obstacle in this task. To address it, this work explores a new solu…

Supervised Video SummarizationVideo Summarization