paper-with-me

Papers

TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use

2025-10-06 · Pengfei He, Zhenwei Dai, Bing He, Hui Liu, Xianfeng Tang, Hanqing Lu, Juanhui Li, Jiayuan Ding, Subhabrata Mukherjee, Suhang Wang, Yue Xing, Jiliang Tang, Benoit Dumoulin arxiv

Large language model (LLM)-based agents increasingly rely on tool use to complete real-world tasks. While existing works evaluate the LLMs' tool use capability, they largely focus on the final answers yet overlook the detailed tool usage trajectory, i.e., whether tools are selected, parameterized, and ordered correctly. We introduce TRAJECT-Bench, a trajectory-aware benchmark to comprehensively evaluate LLMs' tool use capability through diverse tasks with fine-grained evaluation metrics. TRAJECT-Bench pairs high-fidelity, executable tools across practical domains with tasks grounded in production-style APIs, and synthesizes trajectories that vary in breadth (parallel calls) and depth (interdependent chains). Besides final accuracy, TRAJECT-Bench also reports trajectory-level diagnostics, including tool selection and argument correctness, and dependency/order satisfaction. Analyses reveal failure modes such as similar tool confusion and parameter-blind selection, and scaling behavior with tool diversity and trajectory length where the bottleneck of transiting from short to mid-length trajectories is revealed, offering actionable guidance for LLMs' tool use.

📄 PDF Abstract BibTeX arXiv:2510.04550

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents

2026-02-15 · Lingxiang Hu, Yiding Sun, Tianle Xia, Wenwei Li 외 arxiv

While Large Language Model (LLM) agents have made remarkable progress on complex reasoning, evaluating them in real-world environments remains an open problem. Existing benchmarks are largely confined to idealized simula…

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs

2026-06-09 · Joachim Schaeffer, Thomas Jiralerspong, Alexander Panfilov, Guillaume Lajoie 외 arxiv

AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This partially tampers with the untrusted model's trajectory. If the trusted …

Binary Classification

EnvShip-Bench: An Environment-Enhanced Benchmark for Short-Term Vessel Trajectory Prediction

2026-06-13 · Kun Ma, Qilong Han, Chengjing Song, Jingzheng Yao 외 arxiv

Vessel trajectory prediction is important for intelligent shipping, maritime surveillance, and navigation safety. However, existing public maritime AIS resources are often limited by inconsistent forecasting protocols, u…

Trajectory ForecastingTrajectory Prediction

OpenTraj: Assessing Prediction Complexity in Human Trajectories Datasets

2020-10-02 · Javad Amirian, Bingqing Zhang, Francisco Valente Castro, Juan Jose Baldelomar 외

Human Trajectory Prediction (HTP) has gained much momentum in the last years and many solutions have been proposed to solve it. Proper benchmarking being a key issue for comparing methods, this paper addresses the questi…

BenchmarkingPredictionSelf-Driving CarsTrajectory Forecasting+1

CRITERIA: a New Benchmarking Paradigm for Evaluating Trajectory Prediction Models for Autonomous Driving

2023-10-11 · Changhe Chen, Mozhgan PourKeshavarz, Amir Rasouli

Benchmarking is a common method for evaluating trajectory prediction models for autonomous driving. Existing benchmarks rely on datasets, which are biased towards more common scenarios, such as cruising, and distance-bas…

Autonomous DrivingBenchmarkingDiversityPrediction+2