paper-with-me

홈 › Papers

DeliberationBench: When Do More Voices Hurt? A Controlled Study of Multi-LLM Deliberation Protocols

2025-12-14 · Vaarunay Kaushal, Taranveer Singh arxiv

Multi-agent systems where Large Language Models (LLMs) deliberate to form consensus have gained significant attention, yet their practical value over simpler methods remains under-scrutinized. We introduce DELIBERATIONBENCH, a controlled benchmark evaluating three deliberation protocols against a strong baseline of selecting the best response from a pool of model outputs. Across 270 questions and three independent seeds (810 total evaluations), we find a striking negative result: the best-single baseline achieves an 82.5% +- 3.3% win rate, dramatically outperforming the best deliberation protocol(13.8% +- 2.6%). This 6.0x performance gap is statistically significant (p < 0.01) and comes at 1.5-2.5x higher computational cost. Our findings challenge assumptions that complexity enhances quality in multi-LLM systems.

📄 PDF Abstract BibTeX arXiv:2601.08835

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DeliberationBench: A Normative Benchmark for the Influence of Large Language Models on Users' Views

2026-02-22 · Luke Hewitt, Maximilian Kroner Dale, Paul de Font-Reaulx arxiv

As large language models (LLMs) become pervasive as assistants and thought partners, it is important to characterize their persuasive influence on users' beliefs. However, a central challenge is to distinguish "beneficia…

Detoxifying Language Models Risks Marginalizing Minority Voices

2021-04-13 · NAACL 2021 4 · Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan 외

Language models (LMs) must be both safe and equitable to be responsibly deployed in practice. With safety in mind, numerous detoxification techniques (e.g., Dathathri et al. 2020; Krause et al. 2020) have been proposed t…

Text Generation

Who is Authentic Speaker

2024-04-30 · Qiang Huang

Voice conversion (VC) using deep learning technologies can now generate high quality one-to-many voices and thus has been used in some practical application fields, such as entertainment and healthcare. However, voice co…

Speaker RecognitionVoice Conversion

Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024

2024-09-03 · Anmol Guragain, Tianchi Liu, Zihan Pan, Hardik B. Sailor 외

This work details our approach to achieving a leading system with a 1.79% pooled equal error rate (EER) on the evaluation set of the Controlled Singing Voice Deepfake Detection (CtrSVDD). The rapid advancement of generat…

DeepFake DetectionFace SwappingVoice Anti-spoofing

Supervised Fine-tuning with Synthetic Rationale Data Hurts Real-World Disease Prediction

2026-06-09 · Buxin Su, Bingxuan Li, Cheng Qian, Yiwei Wang 외 arxiv

Supervised fine-tuning with synthetic rationale data is widely assumed to improve language model performance on clinical prediction tasks by teaching models not just what to predict but why. We test this assumption on fi…