paper-with-me

홈 › Papers

Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval Evaluation

2026-04-22 · Andrew Klearman, Radu Revutchi, Rohin Garg, Rishav Chakravarti, Samuel Marc Denton, Yuan Xue arxiv

Retrieval quality is the primary bottleneck for accuracy and robustness in retrieval-augmented generation (RAG). Current evaluation relies on heuristically constructed query sets, which introduce a hidden intrinsic bias. We formalize retrieval evaluation as a statistical estimation problem, showing that metric reliability is fundamentally limited by the evaluation-set construction. We further introduce \emph{semantic stratification}, which grounds evaluation in corpus structure by organizing documents into an interpretable global space of entity-based clusters and systematically generating queries for missing strata. This yields (1) formal semantic coverage guarantees across retrieval regimes and (2) interpretable visibility into retrieval failure modes. Experiments across multiple benchmarks and retrieval methods validate our framework. The results expose systematic coverage gaps, identify structural signals that explain variance in retrieval performance, and show that stratified evaluation yields more stable and transparent assessments while supporting more trustworthy decision-making than aggregate metrics.

📄 PDF Abstract BibTeX arXiv:2604.20763

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BiAxisBias: Evaluating LLM Bias Beyond a Single Prompt and a Single Explanation

2026-05-09 · Jialing Gan, Junhao Dong, Songze Li arxiv

LLM bias scores can depend on audit design. We introduce BiAxisBias, a prespecified audit varying task, role, perspective, sentiment, and wording over 200 stereotype statements while retaining forced Selection and Ration…

On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage

2026-07-14 · Vinay Kumar Chaganti arxiv

On-device research agents search a corpus, read sources, and write a cited brief on a personal laptop. Whether their citations are faithful, and at what cost, is unmeasured for a deployable small model. This study fixes …

C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning

2026-05-27 · Yuwei Miao, Gen Li, Yunsheng Zeng, Xiandong Li 외 arxiv

Retrieval-augmented generation combined with reinforcement learning has shown promise for grounding large language models in trustworthy medical evidence. However, existing methods rely on exact-match binary rewards, whi…

Reinforcement Learning

Towards A Fairer Landmark Recognition Dataset

2021-08-19 · Zu Kim, André Araujo, Bingyi Cao, Cam Askew 외

We introduce a new landmark recognition dataset, which is created with a focus on fair worldwide representation. While previous work proposes to collect as many images as possible from web repositories, we instead argue …

Landmark Recognition

Generation-Time vs. Post-hoc Citation: A Holistic Evaluation of LLM Attribution

2025-09-25 · Yash Saxena, Raviteja Bommireddy, Ankur Padia, Manas Gaur arxiv

Trustworthy Large Language Models (LLMs) must cite human-verifiable sources in high-stakes domains such as healthcare, law, academia, and finance, where even small errors can have severe consequences. Practitioners and r…