paper-with-me

Papers

When Does Memory Help Multi-Trajectory Inference for Tool-Use LLM Agents?

2026-05-27 · Xinzhe Li, Yaguang Tao arxiv

Multi-trajectory inference for tool-use LLM agents - generating multiple reasoning attempts and selecting among them - benefits from transferring knowledge across attempts so that later ones avoid the pitfalls of earlier ones. Existing cross-trajectory memory methods (trajectory-level reflection, atomic fact extraction, raw observation injection) are each evaluated under a single inference strategy on a single task, making it unclear whether reported gains reflect properties of the memory abstraction or of the inference method. We propose a unified framework that decomposes memory along two axes -- the scope of transfer (within an expansion vs. across trajectories) and the abstraction of the transferred content -- and evaluate four methods under three inference strategies (best-of-N, beam search, MCTS) on four tool-use benchmarks spanning SQL, knowledge-graph, and CLI environments, in a verifier-free setting that matches the deployment regime of practical agents. The experiment matrix identifies the inference method as a confound: the same memory method produces statistically distinct results under different inference strategies on the same examples. Reflection reaches significance only under MCTS (not under best-of-N); within-expansion injection (conditioning each candidate on prior siblings' outcomes) helps only diversity-starved beam search; and atomic fact extraction is accuracy-neutral but shortens trajectories by 19-26% on tasks with reusable environmental structure.

📄 PDF Abstract BibTeX arXiv:2605.28224

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

When Does Overlap Help? OSU-Mem and a Cell-Conditional Analysis of Trajectory Memory for LLM Agents

2026-06-19 · Mellow Baixuan Chen, Xiangguo Sun arxiv

Long-horizon large language model (LLM) agents accumulate interaction trajectories that quickly exceed any practical prompt budget, and existing memory methods either truncate aggressively and lose non-local evidence or …

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use

2026-05-27 · Abhijit Kumar, Zoey Wu, Mohit Suley arxiv

Humans know when to reach for help e.g. $347 \times 28$ warrants a calculator while $2+2$ does not. Language models do not. Prompt-based approaches can instruct a model when to invoke tools, but this scaffolding does not…

Applying Deep Bidirectional LSTM and Mixture Density Network for Basketball Trajectory Prediction

2017-08-19 · Yu Zhao, Rennong Yang, Guillaume Chevalier, Rajiv Shah 외

Data analytics helps basketball teams to create tactics. However, manual data collection and analytics are costly and ineffective. Therefore, we applied a deep bidirectional long short-term memory (BLSTM) and mixture den…

Time SeriesTime Series AnalysisTrajectory Prediction

Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages

2026-08-14 · Chih-Hsuan Yang, Anjir Ahmed Chowdhury, Cheng-Hau Yang, Weijian Zheng 외 arxiv

Multi-agent reasoning systems often use agreement, confidence, or automated scores to decide which messages should shape a final answer. Such filtering assumes that a message likely to be correct is also worth keeping. Y…

DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models

2026-07-02 · Xi Fang, Weijie Xu, Yingqiang Ge, Yuhui Xu 외 arxiv

Personalization changes what a model says to a user; we show that it can also change the reasoning trajectory used to justify the response. Modern LLMs personalize interactions by storing user attributes, preferences, an…

Instruction Following