paper-with-me

홈 › Papers

Grounded autonomous scrutiny at scale: emergent critique from reproduction of published computational physics papers

2026-04-14 · Haonan Huang arxiv

Autonomous LLM agents now produce complete research artifacts in machine-learning sandboxes, but real computational physics is harder: experiments are first-principles calculations against re-runnable physical ground truth, and meaningful new work almost always builds on a key existing paper. We ask whether such an agent can perform grounded scrutiny of published computational physics - reading a paper, reproducing it from scratch, and surfacing methodological concerns from execution. We deploy a single Claude Opus 4.6 configuration at two complementary scopes. At scale, across 111 open-access Quantum ESPRESSO papers, an autonomous agent runs the read-plan-compute-compare loop and, although never asked to critique, raises substantive methodological concerns on ~42% of papers; 85 of 88 of these critiques (96.6%) surface only after the agent has actually run a calculation, with a reading-only ceiling of 1.8%. Critique emerges from reproduction, not from reading. In depth, on one Nature Communications paper on multiscale device simulation of a 2D-material MOSFET, a fresh agent inheriting a verified reproduction pipeline autonomously produces a 14-concern physics inventory and a complete, submission-form six-page Comment that revises the paper's L_G = 5 nm headline. Two of its L_G = 5 nm headline-challenging attacks - a source-degeneration contact-resistance bound and a Sb-doping degradation ratio - are absent from the published 21-reviewer peer review.

📄 PDF Abstract BibTeX arXiv:2604.12198

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Code Review to Code Critique: Intent, Drift, and Spotlight for AI-Generated Diffs at Scale

2026-07-31 · Chandra Maddila, Mashrur Rashik, Euna Mehnaz Khan, Smriti Jha 외 arxiv

AI coding agents are generating code at volumes that exceed the capacity of traditional peer review. At the same time, existing AI code review tools over-index on low-value suggestions such as style and best practices wh…

EvoIdeator: Evolving Scientific Ideas through Checklist-Grounded Reinforcement Learning

2026-03-23 · Andreas Sauter, Yuyue Zhao, Jacopo Urbani, Wenxiang Hu 외 arxiv

Scientific idea generation is a cornerstone of autonomous knowledge discovery, yet the iterative evolution required to transform initial concepts into high-quality research proposals remains a formidable challenge for La…

Reinforcement Learning

Emergence of Numeric Concepts in Multi-Agent Autonomous Communication

2019-11-04 · Shangmin Guo

With the rapid development of deep learning, most of current state-of-the-art techniques in natural langauge processing are based on deep learning models trained with argescaled static textual corpora. However, we human …

Deep Reinforcement LearningGrounded language learningReinforcement Learning

AesCanvas: A Large-Scale Dataset and Benchmark for Aesthetic Critique and Contextual Suitability

2026-08-27 · Xuanwei Hu, Haoyu Dong, Kejun Wu, Tianyi Liu 외 arxiv

Recent advances in Multimodal Large Language Models (MLLMs) have extended Image Aesthetic Assessment (IAA) beyond scalar scores toward interpretable critique and guidance. Yet existing benchmarks mainly assess intrinsic …

Reply to "Emergent LLM behaviors are observationally equivalent to data leakage"

2025-06-23 · Ariel Flint Ashery, Luca Maria Aiello, Andrea Baronchelli

A potential concern when simulating populations of large language models (LLMs) is data contamination, i.e. the possibility that training data may shape outcomes in unintended ways. While this concern is important and ma…