paper-with-me

홈 › Papers

PRBench: A Standardized Probabilistic Robustness Benchmark

2025-11-03 · Yi Zhang, Zheng Wang, Zhen Chen, Wenjie Ruan, Qing Guo, Siddartha Khastgir, Carsten Maple, Xingyu Zhao arxiv

Deep learning models are notoriously vulnerable to imperceptible perturbations. Most existing research centers on adversarial robustness (AR), which evaluates models under worst-case scenarios by examining the existence of deterministic adversarial examples (AEs). In contrast, probabilistic robustness (PR) adopts a statistical perspective, measuring the probability that predictions remain correct under stochastic perturbations. While PR is widely regarded as a practical complement to AR, dedicated training methods for improving PR are still relatively underexplored, albeit with emerging progress. Among the few PR-targeted training methods, we identify three limitations: i non-comparable evaluation protocols; ii limited comparisons to strong AT baselines despite anecdotal PR gains from AT; and iii no unified framework to compare the generalization of these methods. Thus, we introduce PRBench, the first benchmark dedicated to evaluating improvements in PR achieved by different robustness training methods. PRBench empirically compares most common AT and PR-targeted training methods using a comprehensive set of metrics, including clean accuracy, PR and AR performance, training efficiency, and generalization error (GE). We also provide theoretical analysis on the GE of PR performance across different training methods. Main findings revealed by PRBench include: AT methods are more versatile than PR-targeted training methods in terms of improving both AR and PR performance across diverse hyperparameter settings, while PR-targeted training methods consistently yield lower GE and higher clean accuracy. A leaderboard comprising 229 trained models across 7 datasets and 10 model architectures is publicly available at https://wellzline.github.io/PRBenchLeaderboard/.

📄 PDF Abstract BibTeX arXiv:2511.01724

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial Robustness

Similar Papers 제목 키워드 기반

EPRBench: A High-Quality Benchmark Dataset for Event Stream Based Visual Place Recognition

2026-02-13 · Xiao Wang, Xingxing Xiong, Jinfeng Gao, Xufeng Lou 외 arxiv

Event stream-based Visual Place Recognition (VPR) is an emerging research direction that offers a compelling solution to the instability of conventional visible-light cameras under challenging conditions such as low illu…

Visual Place RecognitionRepresentation Learning

PRBench: End-to-end Paper Reproduction in Physics Research

2026-03-29 · Shi Qiu, Junyi Deng, Yiwei Deng, Haoran Dong 외 arxiv

AI agents powered by large language models exhibit strong reasoning and problem-solving capabilities, enabling them to assist scientific research tasks such as formula derivation and code generation. However, whether the…

Code Generation

SWE-PRBench: Benchmarking AI Code Review Quality Against Pull Request Feedback

2026-03-27 · Deepak Kumar arxiv

We introduce SWE-PRBench, a benchmark of 350 pull requests with human-annotated ground truth for evaluating AI code review quality. Evaluated against an LLM-as-judge framework validated at kappa=0.75, 8 frontier models d…

Code Generation

PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning

2025-11-14 · Afra Feyza Akyürek, Advait Gosai, Chen Bo Calvin Zhang, Vipul Gupta 외 arxiv

Frontier model progress is often measured by academic benchmarks, which offer a limited view of performance in real-world professional contexts. Existing evaluations often fail to assess open-ended, economically conseque…

AutoPR: Let's Automate Your Academic Promotion!

2025-10-10 · Qiguang Chen, Zheng Yan, Mingda Yang, Libo Qin 외 arxiv

As the volume of peer-reviewed research surges, scholars increasingly rely on social platforms for discovery, while authors invest considerable effort in promoting their work to ensure visibility and citations. To stream…