paper-with-me

Papers

Beyond Consensus: Downward Bias and Role Asymmetry in Multi-Agent LLM Judges for Subjective Evaluation

2026-08-31 · Minsoo Song, Chanwoo Kim, Sugyeong Eo, Chanjun Park arxiv

Multi-Agent Debate (MAD) has been widely adopted to improve LLM-based evaluation by prompting multiple agents to negotiate and reach a consensus. However, for subjective rubric-based scoring, inter-agent agreement does not guarantee alignment with human judgments. In this paper, we compare a single-judge baseline against a consensus-based MAD protocol on subjective evaluation tasks and design three ablations to isolate the impact of role prompting, multi-round interaction, and explicit score sharing. Evaluations across six LLMs show that the single-judge baseline achieves the strongest human alignment on average across six judge models, whereas MAD shows degradation in human alignment on both tasks. Our ablations demonstrate that this performance drop stems primarily from asymmetric role prompting rather than the interaction itself. Specifically, assigning a strict judge role introduces a systematic downward bias that the consensus process fails to correct. The central finding is that this bias reflects strict-stance dominance beyond averaging: the consensus score falls well beyond the arithmetic midpoint of the standalone strict and lenient conditions, rather than averaging them out. Removing role asymmetry (Symmetric MAD) largely recovers baseline performance, while masking peer scores widens inter-agent disagreement on average and worsens average human alignment. These findings demonstrate that multi-agent consensus can enforce artificial agreement at the expense of true human alignment, revealing a structural limitation in consensus-style, role-specialized MAD protocols for subjective scoring.

📄 PDF Abstract BibTeX arXiv:2608.30373

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment

2026-04-21 · Bobo Li, Rui Wu, Zibo Ji, Meishan Zhang 외 arxiv

Large Language Model agents have rapidly evolved from static text generators into dynamic systems capable of executing complex autonomous workflows. To enhance reliability, multi-agent frameworks assigning specialized ro…

Bridging the RGB-IR Gap: Consensus and Discrepancy Modeling for Text-Guided Multispectral Detection

2026-04-13 · Jiaqi Wu, Zhen Wang, Enhao Huang, Kangqing Shen 외 arxiv

Text-guided multispectral object detection uses text semantics to guide semantic-aware cross-modal interaction between RGB and IR for more robust perception. However, notable limitations remain: (1) existing methods ofte…

Multispectral Object Detection

Generalized Autoregressive Score asymmetric Laplace Distribution and Extreme Downward Risk Prediction

2020-08-04 · Hong Shaopeng

Due to the skessed distribution, high peak and thick tail and asymmetry of financial return data, it is difficult to describe the traditional distribution. In recent years, generalized autoregressive score (GAS) has been…

SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models

2026-04-16 · Binxian Su, Haoye Lou, Shucheng Zhu, Weikang Wang 외 arxiv

Large language models (LLMs) are being increasingly used in urban planning, but since gendered space theory highlights how gender hierarchies are embedded in spatial organization, there is concern that LLMs may reproduce…

Story Generation

To Asymmetry and Beyond: Structured Pruning of Sequence to Sequence Models for Improved Inference Efficiency

2023-04-05 · Daniel Campos, ChengXiang Zhai

Sequence-to-sequence language models can be used to produce abstractive summaries which are coherent, relevant, and concise. Still, model sizes can make deployment in latency-sensitive or web-scale implementations diffic…

Decoder