paper-with-me

홈 › Papers

SGR-Bench: Benchmarking Search Agents on State-Gated Retrieval

2026-05-21 · Ningyuan Li, Haiyang Shen, Mugeng Liu, Yudong Han, Zhuofan Shi, Sixiong Xie, Yun Ma arxiv

Recent advances in large language models and tool-using agents have expanded the range of benchmarked web tasks. Yet an important class of specialized retrieval tasks remains undercharacterized. On many specialized data-retrieval websites, answer-bearing evidence becomes accessible only after establishing the correct site-specific retrieval state through filters, views, hierarchies, or scopes. We term this capability state-gated retrieval (SGR). We introduce SGR-Bench, a benchmark for this setting containing 100 expert-curated tasks spanning six source families and 12 public data ecosystems. Each task requires discovering the appropriate website and configuring its site-specific retrieval state to produce a structured answer. SGR-Bench pairs constraint-guided and goal-oriented formulations of the same underlying problems, enabling controlled comparisons between explicit and implicit guidance for state-gated retrieval. We evaluate eight CLI-based agentic LLM systems and three commercial search-agent products. On SGR-Bench, the strongest system reaches only 66.18% item-level F1, while row-level F1 remains much lower. A manual audit of 156 analyzable failed CLI trajectories shows why: agents often reach a relevant web source, but establish the wrong site-specific retrieval state. Retrieval-scope drift (37.2%) and criterion mismatch (27.6%) dominate, whereas final answer composition accounts for only 10.3%. The dataset and single-case evaluation instructions are available at https://huggingface.co/datasets/PKUAIWeb/SGR-BENCH.

📄 PDF Abstract BibTeX arXiv:2605.22219

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Delegated Authorization for Agents Constrained to Semantic Task-to-Scope Matching

2025-10-30 · Majed El Helou, Chiara Troiani, Benjamin Ryder, Jean Diaconu 외 arxiv

Authorizing Large Language Model driven agents to dynamically invoke tools and access protected resources introduces significant risks, since current methods for delegating authorization grant overly broad permissions an…

EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent

2026-06-16 · Zeyao Du, Tong Li, Yanci Zhang, Haibo Zhang arxiv

As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly in the query, recorded in a profile, or revealed only when the right question is a…

CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios

2026-05-08 · Taein Lim, Seongyong Ju, Munhyeok Kim, Hyunjun Kim 외 arxiv

Large language models (LLMs) are increasingly deployed as autonomous agents in offensive cybersecurity. In this paper, we reveal an interesting phenomenon: different agents exhibit distinct attack patterns. Specifically,…

Towards a Playground to Democratize Experimentation and Benchmarking of AI Agents for Network Troubleshooting

2025-07-01 · Zhihao Wang, Alessandro Cornacchia, Franco Galante, Carlo Centofanti 외 arxiv

Recent research has demonstrated the effectiveness of Artificial Intelligence (AI), and more specifically, Large Language Models (LLMs), in supporting network configuration synthesis and automating network diagnosis task…

FML-bench: Benchmarking Machine Learning Agents for Scientific Research

2025-10-12 · Qiran Zou, Hou Hei Lam, Wenhao Zhao, Yiming Tang 외 arxiv

Large language models (LLMs) have sparked growing interest in machine learning research agents that can autonomously propose ideas and conduct experiments. However, existing benchmarks predominantly adopt an engineering-…