paper-with-me

홈 › Papers

Beyond Scaffold Splits: Structural-Frontier Evaluation Reveals Hidden Failures in ADMET Models

2026-07-12 · Jiacheng Zheng, Chang Guo, Zixuan Wang, Xinyu Liu arxiv

Molecular property models are commonly evaluated by holding out Bemis-Murcko scaffolds, yet a scaffold identifier is only one notion of chemical unfamiliarity. We introduce a label-free structural-frontier split that reserves the sparsest and most physicochemically remote scaffold groups, and evaluate it on six public experimental or curated ADMET tasks. Against a 70/10/20 scaffold control with identical acyclic grouping, the frontier inflates equally weighted primary error with a taskwise median of 87.0% and a skew-sensitive mean of 130.3% (descriptive task/seed bootstrap interval, 52.1-246.0%). The mean falls to 75.9% once BBB is removed; that endpoint is the one whose score ranking inverts at the frontier. A message-passing graph-network control still shows a large gap (mean 82.8% over four tasks) and does not invert, so a low-capacity head does not explain the effect. We also test Multi-View Frontier Risk Extrapolation (MV-FREX), a count-adjusted tail-risk penalty over four molecular views, and treat it as a falsifiable probe. It changes normalized frontier error by only 0.16% relative to empirical risk minimization for the perceptron head (interval, -0.43-0.84%) and by -1.9% for the graph network; three fixed robust-penalty controls are likewise inconclusive. Against the published Lo-Hi and DataSAIL splitters, the frontier inflates error more on average, though no split is uniformly hardest. An audit of 31,561 marine natural products further shows that OOD status and agreement with legacy ADMET predictions depend on the molecular view, endpoint, and teacher coverage. Split construction and label provenance are important evaluation constraints in their own right, and the tested training penalties do not resolve the frontier failures we observe.

📄 PDF Abstract BibTeX arXiv:2607.10729

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scaffold Splits Overestimate Virtual Screening Performance

2024-06-02 · Qianrong Guo, Saiveth Hernandez-Hernandez, Pedro J Ballester

Virtual Screening (VS) of vast compound libraries guided by Artificial Intelligence (AI) models is a highly productive approach to early drug discovery. Data splitting is crucial for better benchmarking of such AI models…

BenchmarkingClusteringDrug DiscoveryMolecular Property Prediction+1

How Far Can You Grow? Characterizing the Extrapolation Frontier of Graph Generative Models for Materials Science

2026-02-10 · Can Polat, Erchin Serpedin, Mustafa Kurban, Hasan Kurban arxiv

Every generative model for crystalline materials harbors a critical structure size beyond which its outputs become unreliable; we call this the extrapolation frontier. Despite its consequences for nanomaterial design, th…

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

2026-03-08 · David Gringras arxiv

A safety score earned on a benchmark need not predict how the same model behaves once it is wrapped in an agentic scaffold the benchmark never tested. We ran six frontier models through four deployment configurations (di…

Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead

2026-09-15 · Fengshuo Liu, Ying Liu, Ruize Sun, Lie Luo 외 arxiv

Small differences on coding-agent leaderboards are often read as an ordering of systems. We audit whether the published verdicts support this reading, using 254 SWE-bench submissions across four splits without running mo…

PieArena: Ranking and Profiling Language Agents in Realistic Negotiation Scenarios

2026-02-05 · Chris Zhu, Sasha Cui, Will Sanok Dufallo, Runzhi Jin 외 arxiv

We present an in-depth evaluation of LLMs' ability to negotiate, a central business task requiring strategic reasoning, theory of mind, and economic value creation. To do so, we introduce PieArena, a large-scale negotiat…