paper-with-me

홈 › Papers

Whose Name Comes Up? II: Benchmarking and Intervention-Based Auditing of LLM-Based Scholar Recommendation

2026-02-09 · Lisette Espín-Noboa, Gonzalo Gabriel Méndez arxiv

Large language models (LLMs) are now used for academic expert recommendation. Existing audits typically evaluate such recommendations in isolation, ignoring end-user inference-time interventions. Thus, it remains unclear whether failures (e.g., refusals, hallucinations, uneven coverage) stem from model choice or deployment decisions. We introduce LLMScholarBench, a benchmark for auditing LLM-based scholar recommendation that jointly evaluates model infrastructure and end-user interventions across multiple tasks. LLMScholarBench measures technical quality and social representation using nine metrics. We instantiate the benchmark in physics expert recommendation and audit 22 LLMs under temperature variation, representation-constrained prompting, and retrieval-augmented generation (RAG) via web search. Our results show that each intervention entails distinct tradeoffs. Higher temperature degrades validity, consistency, and factuality. Representation-constrained prompting improves diversity at the expense of factuality, while RAG primarily improves technical quality while reducing diversity and parity. Overall, end-user interventions reshape trade-offs rather than providing uniform gains. LLMScholarBench makes all these dynamics auditable across models and interventions in LLM-based scholar recommendations.

📄 PDF Abstract BibTeX arXiv:2602.08873

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Quantifying Model Uniqueness in Heterogeneous AI Ecosystems

2026-01-30 · Lei You arxiv

As AI systems evolve from isolated predictors into complex, heterogeneous ecosystems of foundation models and specialized adapters, distinguishing genuine behavioral novelty from functional redundancy becomes a critical …

Smiling Women Pitching Down: Auditing Representational and Presentational Gender Biases in Image Generative AI

2023-05-17 · Luhang Sun, Mian Wei, Yibing Sun, Yoo Ji Suh 외

Generative AI models like DALL-E 2 can interpret textual prompts and generate high-quality images exhibiting human creativity. Though public enthusiasm is booming, systematic auditing of potential gender biases in AI-gen…

Benchmarking

RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning Models

2026-02-19 · Yunseok Han, Yejoon Lee, Jaeyoung Do arxiv

Large Reasoning Models (LRMs) exhibit strong performance, yet often produce rationales that sound plausible but fail to reflect their true decision process, undermining reliability and trust. We introduce a formal framew…

Auditing Near-Optimal Policies Can Be Exponentially Hard: Conditional Query Lower Bounds via Occupancy Rashomon Capacity

2026-05-29 · Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma arxiv

When many reinforcement-learning policies achieve near-optimal return, a post-hoc auditor may have to distinguish among many behaviorally distinct but return-equivalent policies. We formalize this phenomenon through an o…

An Auditing Test To Detect Behavioral Shift in Language Models

2024-10-25 · Leo Richter, Xuanli He, Pasquale Minervini, Matt J. Kusner

As language models (LMs) approach human-level performance, a comprehensive understanding of their behavior becomes crucial. This includes evaluating capabilities, biases, task performance, and alignment with societal val…

BenchmarkingChange DetectionRed Teaming