paper-with-me

홈 › Papers

Measuring Collapse and Correction in Homogeneous-Panel LLM Debate

2026-09-28 · Xin Li, Mengbing Liu, Chau Yuen hf

Multi-agent large language model (LLM) debate is often evaluated by whether final answers improve, but movement is not necessarily improvement: the same discussion can rescue an initially wrong majority or destroy an initially correct one. Standard final-accuracy evaluations conflate these opposing mechanisms. We introduce an auditable protocol for homogeneous debate on multiple-choice questions (MCQs) that records each run as a transition ledger over collapse, correction, onset, and signed intervention utility. On 6,925 MMLU-Pro debates, the protocol identifies 253 collapses and a parallel correction ledger that changes how interventions should be judged. Replay experiments reveal the central tradeoff: a leave-one-model-out probe-gated freeze prevents 29 collapses but loses 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy. A compact pre-debate 8-probe screen is a triage signal: its unadjusted family-level association with conditional-collapse risk is high (G=7, Spearman rho=0.893, exact two-sided p=0.0123), but initial-majority accuracy is a close comparator (rho=0.821; family partial rho=0.767, p=0.0877), so we do not treat it as calibrated or capability-adjusted prediction. Round-level traces localize many collapses to the first debate round, where early disagreement can precede both harmful cascades and useful recovery. We release replayable schemas, coders, audits, cost cards, and zero-API rebuild scripts so future model-scaffold rows can be compared under the same denominators and signed utility ledger.

📄 PDF Abstract BibTeX arXiv:2609.35279

Code (1)

Valiant-Cat/hfpaper

Similar Papers 제목 키워드 기반

The Cost of Consensus: Isolated Self-Correction Prevails Over Unguided Homogeneous Multi-Agent Debate

2026-04-29 · Blaž Bertalanič, Carolina Fortuna arxiv

Multi-agent debate, where teams of LLMs iteratively exchange rationales and vote on answers, is widely deployed under the assumption that peer review filters hallucinations. Yet the failure dynamics of homogeneous debate…

The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size

2026-05-31 · Blaž Bertalanič, Carolina Fortuna arxiv

Inference-time multi-agent LLM scaling lacks a shared unit: counting nominal agents conflates cost with independent evidence. We derive a two-parameter scaling law $R(N) = N_\text{eff}/N = 1/(1+c(N-1)N^{-β})$ where the r…

Novel Feature-Based Clustering of Micro-Panel Data (CluMP)

2018-07-16 · Lukas Sobisek, Maria Stachova, Jan Fojtik

Micro-panel data are collected and analysed in many research and industry areas. Cluster analysis of micro-panel data is an unsupervised learning exploratory method identifying subgroup clusters in a data set which inclu…

Clustering

Debatrix: Multi-dimensional Debate Judge with Iterative Chronological Analysis Based on LLM

2024-03-12 · Jingcong Liang, Rong Ye, Meng Han, Ruofei Lai 외

How can we construct an automated debate judge to evaluate an extensive, vibrant, multi-turn debate? This task is challenging, as judging a debate involves grappling with lengthy texts, intricate argument relationships, …

Overidentification in Shift-Share Designs

2024-04-25 · Jinyong Hahn, Guido Kuersteiner, Andres Santos, Wavid Willigrod

This paper studies the testability of identifying restrictions commonly employed to assign a causal interpretation to two stage least squares (TSLS) estimators based on Bartik instruments. For homogeneous effects models …

valid