paper-with-me

Papers

PATHWAYS: Evaluating Investigation and Context Discovery in AI Web Agents

2026-02-05 · Shifat E. Arman, Syed Nazmus Sakib, Tapodhir Karmakar Taton, Nafiul Haque, Shahrear Bin Amin arxiv

We introduce PATHWAYS, a benchmark of 250 multi-step decision tasks that test whether web-based agents can discover and correctly use hidden contextual information. Across both closed and open models, agents typically navigate to relevant pages but retrieve decisive hidden evidence in only a small fraction of cases. When tasks require overturning misleading surface-level signals, performance drops sharply to near chance accuracy. Agents frequently hallucinate investigative reasoning by claiming to rely on evidence they never accessed. Even when correct context is discovered, agents often fail to integrate it into their final decision. Providing more explicit instructions improves context discovery but often reduces overall accuracy, revealing a tradeoff between procedural compliance and effective judgement. Together, these results show that current web agent architectures lack reliable mechanisms for adaptive investigation, evidence integration, and judgement override.

📄 PDF Abstract BibTeX arXiv:2602.05354

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents

2026-04-13 · Daniel Begimher, Cristian Leo, Jack Huang, Pat Gaw 외 arxiv

We present SIR-Bench, a benchmark of 794 test cases for evaluating autonomous security incident response agents that distinguishes genuine forensic investigation from alert parroting. Derived from 129 anonymized incident…

Evaluating Memory Condensation Strategies for Coding Agents in Data-Driven Scientific Discovery

2026-05-13 · Renuka Chintalapati, Sid Raskar, Anurag Acharya, Jared Willard 외 arxiv

Coding agents accumulate extensive context during long-running tasks, yet fixed context windows force practitioners to choose between truncation and task failure. While numerous memory condensation strategies have been p…

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?

2026-06-23 · Yuru Wang, Lejun Cheng, Yuxin Zuo, Sihang Zeng 외 arxiv

We introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publications, designed to evaluate whether AI coding agents can move beyond reproduction toward discovery on r…

Separable Pathways for Causal Reasoning: How Architectural Scaffolding Enables Hypothesis-Space Restructuring in LLM Agents

2026-04-21 · John Alderete, Sebastian Benthal, Connie Xu, John Xing arxiv

Causal discovery through experimentation and intervention is fundamental to robust problem solving. It requires not just updating beliefs within a fixed framework but revising the hypothesis space itself, a capacity curr…

DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents

2024-06-10 · Peter Jansen, Marc-Alexandre Côté, Tushar Khot, Erin Bransom 외

Automated scientific discovery promises to accelerate progress across scientific domains. However, developing and evaluating an AI agent's capacity for end-to-end scientific reasoning is challenging as running real-world…

Benchmarkingscientific discovery