paper-with-me

Papers

PeerRank: Autonomous LLM Evaluation Through Web-Grounded, Bias-Controlled Peer Review

2026-02-01 · Yanki Margalit, Erni Avram, Ran Taig, Oded Margalit, Nurit Cohen-Inger arxiv

Evaluating large language models typically relies on human-authored benchmarks, reference answers, and human or single-model judgments, approaches that scale poorly, become quickly outdated, and mismatch open-world deployments that depend on web retrieval and synthesis. We introduce PeerRank, a fully autonomous end-to-end evaluation framework in which models generate evaluation tasks, answer them with category-scoped live web grounding, judge peer responses and aggregate dense peer assessments into relative performance estimates, without human supervision or gold references. PeerRank treats evaluation as a multi-agent process where each model participates symmetrically as task designer, respondent, and evaluator, while removing biased judgments. In a large-scale study over 12 commercially available models and 420 autonomously generated questions, PeerRank produces stable, discriminative rankings and reveals measurable identity and presentation biases. Rankings are robust, and mean peer scores agree with Elo. We further validate PeerRank on TruthfulQA and GSM8K, where peer scores correlate with objective accuracy. Together, these results suggest that bias-aware peer evaluation with selective web-grounded answering can scale open-world LLM assessment beyond static and human curated benchmarks.

📄 PDF Abstract BibTeX arXiv:2602.02589

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The PeerRank Method for Peer Assessment

2014-05-28 · Toby Walsh

We propose the PeerRank method for peer assessment. This constructs a grade for an agent based on the grades proposed by the agents evaluating the agent. Since the grade of an agent is a measure of their ability to grade…

AfriStereo: A Culturally Grounded Dataset for Evaluating Stereotypical Bias in Large Language Models

2025-11-27 · Yann Le Beux, Oluchi Audu, Oche D. Ankeli, Dhananjay Balakrishnan 외 arxiv

Existing AI bias evaluation benchmarks largely reflect Western perspectives, leaving African contexts underrepresented and enabling harmful stereotypes in applications across various domains. To address this gap, we intr…

Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement

2026-06-16 · Ramaravind Kommiya Mothilal, Terry Jingchen Zhang, Raiyan Ahmed, Zhijing Jin 외 arxiv

Evaluations of social bias in LLMs largely focus on whether models generate or imply biased content. However, as LLMs are increasingly used as judges of bias, they may exhibit social biases in subtler ways in how they ev…

Logical Reasoning

Trustworthy AI: Ensuring Reliability and Accountability from Models to Agents

2026-05-09 · Carol Xuan Long arxiv

In this thesis, we develop algorithms with theoretical guarantees for ensuring reliability and accountability of Machine Learning (ML) systems. As ML systems evolve from predictive models to generative models and autonom…

SB-Bench: Stereotype Bias Benchmark for Large Multimodal Models

2025-02-12 · Vishal Narnaware, Ashmal Vayani, Rohit Gupta, Swetha Sirnam 외

Stereotype biases in Large Multimodal Models (LMMs) perpetuate harmful societal prejudices, undermining the fairness and equity of AI applications. As LMMs grow increasingly influential, addressing and mitigating inheren…

FairnessMultiple-choice