paper-with-me

홈 › Papers

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks

2025-07-23 · Linbo Cao, Jinman Zhao arxiv

As frontier language models increasingly saturate standard QA benchmarks, concerns about data contamination, memorization, and escalating dataset creation costs persist. We propose a debate-driven evaluation paradigm that transforms any existing QA dataset into structured adversarial debates--where one model is given the official answer to defend, and another constructs and defends an alternative answer--adjudicated by a judge model blind to the correct solution. By forcing multi-round argumentation, this approach substantially increases difficulty while penalizing shallow memorization, yet reuses QA items to reduce curation overhead. We make two main contributions: (1) an evaluation pipeline to systematically convert QA tasks into debate-based assessments, and (2) a public benchmark that demonstrates our paradigm's effectiveness on a subset of MMLU-Pro questions, complete with standardized protocols and reference models. Empirical results validate the robustness of the method and its effectiveness against data contamination--a Llama 3.1 model fine-tuned on test questions showed dramatic accuracy improvements (50% -> 82%) but performed worse in debates. Results also show that even weaker judges can reliably differentiate stronger debaters, highlighting how debate-based evaluation can scale to future, more capable systems while maintaining a fraction of the cost of creating new benchmarks. Overall, our framework underscores that "pretraining on the test set is no longer all you need," offering a sustainable path for measuring the genuine reasoning ability of advanced language models.

📄 PDF Abstract BibTeX arXiv:2507.17747

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Stay Focused: Problem Drift in Multi-Agent Debate

2025-02-26 · Jonas Becker, Lars Benedikt Kaesberg, Andreas Stephan, Jan Philip Wahle 외

Multi-agent debate - multiple instances of large language models discussing problems in turn-based interaction - has shown promise for solving knowledge and reasoning tasks. However, these methods show limitations when s…

Instruction Following

ArgCMV: An Argument Summarization Benchmark for the LLM-era

2025-08-27 · Omkar Gurjar, Agam Goyal, Eshwar Chandrasekharan arxiv

Key point extraction is an important task in argument summarization which involves extracting high-level short summaries from arguments. Existing approaches for KP extraction have been mostly evaluated on the popular Arg…

Q-STRUM Debate: Query-Driven Contrastive Summarization for Recommendation Comparison

2025-02-18 · George-Kirollos Saad, Scott Sanner

Query-driven recommendation with unknown items poses a challenge for users to understand why certain items are appropriate for their needs. Query-driven Contrastive Summarization (QCS) is a methodology designed to addres…

Evaluating LLM-Driven Summarisation of Parliamentary Debates with Computational Argumentation

2026-04-21 · Eoghan Cunningham, Derek Greene, James Cross, Antonio Rago arxiv

Understanding how policy is debated and justified in parliament is a fundamental aspect of the democratic process. However, the volume and complexity of such debates mean that outside audiences struggle to engage. Meanwh…

An LLM-Driven Multi-Agent Debate System for Mendelian Diseases

2025-04-10 · Xinyang Zhou, Yongyong Ren, Qianqian Zhao, Daoyi Huang 외

Accurate diagnosis of Mendelian diseases is crucial for precision therapy and assistance in preimplantation genetic diagnosis. However, existing methods often fall short of clinical standards or depend on extensive datas…

DiagnosticLanguage ModelingLanguage Modelling