paper-with-me

Papers

Beyond pass@1: A Reliability Science Framework for Long-Horizon LLM Agents

2026-03-31 · Aaditya Khanal, Yangyang Tao, Junxiu Zhou arxiv

Existing benchmarks measure capability -- whether a model succeeds on a single attempt -- but production deployments require reliability -- consistent success across repeated attempts on tasks of varying duration. We show these properties diverge systematically as task duration grows, and that pass@1 on short tasks is structurally blind to this divergence. We introduce a reliability science framework for long-horizon LLM agents with four metrics: Reliability Decay Curve (RDC), Variance Amplification Factor (VAF), Graceful Degradation Score (GDS), and Meltdown Onset Point (MOP). We evaluate 10 models across 23,392 episodes on a 396-task benchmark spanning four duration buckets and three domains. Key findings: (1) reliability decay is domain-stratified -- SE GDS drops from 0.90 to 0.44 while document processing is nearly flat (0.74 to 0.71); (2) VAF bifurcates by capability tier -- high VAF is a capability signature, not an instability signal; (3) capability and reliability rankings diverge substantially, with multi-rank inversions at long horizons; (4) frontier models have the highest meltdown rates (up to 19%) because they attempt ambitious multi-step strategies that sometimes spiral; and (5) memory scaffolds universally hurt long-horizon performance across all 10 models. These results motivate reliability as a first-class evaluation dimension alongside capability.

📄 PDF Abstract BibTeX arXiv:2603.29231

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

What Is Actually Being Annotated? Inter-Prompt Reliability as a Measurement Problem in LLM-Based Social Science Labeling

2026-04-02 · Jingyuan Liu arxiv

Large language models (LLMs) are increasingly used for annotation in computational social science, yet their methodological reliability under prompt variation remains unclear. This paper introduces Inter-Prompt Reliabili…

Automated Annotation of Evolving Corpora for Augmenting Longitudinal Network Data: A Framework Integrating Large Language Models and Expert Knowledge

2025-03-03 · Xiao Liu, Zirui Wu, Jiayi Li, Zhicheng Shao 외

Longitudinal network data are essential for analyzing political, economic, and social systems and processes. In political science, these datasets are often generated through human annotation or supervised machine learnin…

Thinking Outside the [Chat]Box: Bridging Computer Science and Industrial Design for Cognitive-Inclusive Generative AI

2026-06-12 · Virginia Francisco, Daniel Guasch, Raquel Hervás arxiv

Current Generative AI (GenAI) interfaces remain largely constrained to chatbox interaction, which can impose high cognitive demands on users and create substantial barriers for people with intellectual disabilities (ID),…

AI-for-Science Low-code Platform with Bayesian Adversarial Multi-Agent Framework

2026-03-03 · Zihang Zeng, Jiaquan Zhang, Pengze Li, Yuan Qi 외 arxiv

Large Language Models (LLMs) demonstrate potentials for automating scientific code generation but face challenges in reliability, error propagation in multi-agent workflows, and evaluation in domains with ill-defined suc…

Prompt EngineeringCode Generation

Depth and Autonomy: A Framework for Evaluating LLM Applications in Social Science Research

2025-10-29 · Ali Sanaei, Ali Rajabzadeh arxiv

Large language models (LLMs) are increasingly utilized by researchers across a wide range of domains, and qualitative social science is no exception; however, this adoption faces persistent challenges, including interpre…