paper-with-me

홈 › Papers

SWE-PRBench: Benchmarking AI Code Review Quality Against Pull Request Feedback

2026-03-27 · Deepak Kumar arxiv

We introduce SWE-PRBench, a benchmark of 350 pull requests with human-annotated ground truth for evaluating AI code review quality. Evaluated against an LLM-as-judge framework validated at kappa=0.75, 8 frontier models detect only 15-31% of human-flagged issues on the diff-only configuration, demonstrating that AI code review remains far below human expert performance despite strong results on code generation benchmarks. Pull requests are drawn from active open-source repositories, filtered from 700 candidates using a Repository Quality Score, and evaluated under three frozen context configurations: diff only (config_A), diff with file content (config_B), and full context (config_C), enabling systematic ablation of context provision strategies. All 8 models degrade monotonically from config_A to config_C, even when context is provided via structured semantic layers including AST-extracted function context and import graph resolution. The dominant mechanism is a collapse of Type2_Contextual issue detection at config_B, consistent with attention dilution in long contexts: a structured 2,000-token diff-with-summary prompt outperforms a 2,500-token full-context prompt enriched with execution context, behaviour mapping, and test signatures across all 8 models. The top four models are statistically indistinguishable (mean score 0.147-0.153) while a clear tier gap separates them from the remaining four (mean score <= 0.113). Dataset, contexts, annotations, and evaluation harness are released publicly.

📄 PDF Abstract BibTeX arXiv:2603.26130

Code (0)

등록된 구현이 없습니다.

Tasks

Code Generation

Similar Papers 제목 키워드 기반

EPRBench: A High-Quality Benchmark Dataset for Event Stream Based Visual Place Recognition

2026-02-13 · Xiao Wang, Xingxing Xiong, Jinfeng Gao, Xufeng Lou 외 arxiv

Event stream-based Visual Place Recognition (VPR) is an emerging research direction that offers a compelling solution to the instability of conventional visible-light cameras under challenging conditions such as low illu…

Visual Place RecognitionRepresentation Learning

AutoPR: Let's Automate Your Academic Promotion!

2025-10-10 · Qiguang Chen, Zheng Yan, Mingda Yang, Libo Qin 외 arxiv

As the volume of peer-reviewed research surges, scholars increasingly rely on social platforms for discovery, while authors invest considerable effort in promoting their work to ensure visibility and citations. To stream…

PRBench: End-to-end Paper Reproduction in Physics Research

2026-03-29 · Shi Qiu, Junyi Deng, Yiwei Deng, Haoran Dong 외 arxiv

AI agents powered by large language models exhibit strong reasoning and problem-solving capabilities, enabling them to assist scientific research tasks such as formula derivation and code generation. However, whether the…

Code Generation

IPR: Intelligent Prompt Routing with User-Controlled Quality-Cost Trade-offs

2025-09-08 · Aosong Feng, Balasubramaniam Srinivasan, Yun Zhou, Zhichao Xu 외 arxiv

Routing incoming queries to the most cost-effective LLM while maintaining response quality poses a fundamental challenge in optimizing performance-cost trade-offs for large-scale commercial systems. We present IPR\, -- \…

PRBench: A Standardized Probabilistic Robustness Benchmark

2025-11-03 · Yi Zhang, Zheng Wang, Zhen Chen, Wenjie Ruan 외 arxiv

Deep learning models are notoriously vulnerable to imperceptible perturbations. Most existing research centers on adversarial robustness (AR), which evaluates models under worst-case scenarios by examining the existence …

Adversarial Robustness