paper-with-me

홈 › Papers

Stress-testing large language model agents in a robotic chemistry laboratory

2026-07-25 · Lulu Guo, Yingkai Sun, Xiaobo Li, Luyao Ge, Ziming Wang, Haitao Zheng, Jingyu Li, Huijuan Zhang, Bingxu Chen, Daobin Liu, Yuebo Liu, Jie Li, Xiaohui Li, Linjiang Chen, Yi Luo, Jun Jiang arxiv

AI is evaluated through knowledge, reasoning and plan generation, yet scientific agency requires reliable physical action and adaptation to evidence. Here, we use a robotic chemistry laboratory as a physical-world testbed to make scientific agency measurable. Its 45 modular workstations exposed as machine-readable skills enabled 4,608 trials. Only 3.3% of trials produced expert-assessed executable workflows under laboratory constraints; even the best system achieved 28.1%. Long-horizon planning remained a challenge: only three executable workflows exceeded 30 operations, although the longest contained 44. Across five rounds, experimental feedback prompted local adjustments but no workflow-level replanning or analytical-method redesign. By making physical executability and evidence-driven replanning measurable, our study provides an evidence-based assessment of deployment readiness and a diagnostic framework to guide closed-loop improvements towards physically grounded autonomous research.

📄 PDF Abstract BibTeX arXiv:2607.23045

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Stress Testing Concept Erasure with Large Language Model Agents

2026-07-20 · Yuyang Xue, Feng Chen, Zhihua Liu, Edward Moroshko 외 arxiv

Concept erasure aims to remove semantic concepts from a trained generative model and is increasingly important for responsible AI deployment. However, verifying whether a model has robustly removed targeted concepts rema…

RobotArena $\infty$: Scalable Robot Benchmarking via Real-to-Sim Translation

2025-10-27 · Yash Jangir, Yidi Zhang, Pang-Chi Lo, Kashu Yamazaki 외 arxiv

The pursuit of robot generalists, agents capable of performing diverse tasks across diverse environments, demands rigorous and scalable evaluation. Yet real-world testing of robot policies remains fundamentally constrain…

Robot Manipulation

StressWeb: A Diagnostic Benchmark for Web Agent Robustness under Realistic Interaction Variability

2026-03-27 · Haoyue Bai, Dong Wang, Long Chen, Bingguang Hao 외 arxiv

Large language model-based web agents have demonstrated strong performance on realistic web interaction tasks. However, existing evaluations are predominantly conducted under relatively stable and well-behaved interactio…

IntenTest: Stress Testing for Intent Integrity in API-Calling LLM Agents

2025-06-09 · Shiwei Feng, Xiangzhe Xu, Xuan Chen, Kaiyuan Zhang 외

LLM agents are increasingly deployed to automate real-world tasks by invoking APIs through natural language instructions. While powerful, they often suffer from misinterpretation of user intent, leading to the agent's ac…

software testing

Stress-testing medical large language models reveals latent safety pathology beyond benchmark accuracy

2026-06-06 · Yuan Shen, Xiaojun Wu, Linghua Yu arxiv

Large language models (LLMs) are entering clinical practice based on benchmark accuracy that may fail to detect safety-relevant failure modes. Here we present AI-MASLD, a stress-audit framework that adapts the logic of m…

Information Extraction