paper-with-me

홈 › Papers

What Do Agents Learn from Trajectory-SFT: Semantics or Interfaces?

2026-02-02 · Weizheng Gu, Chengze Li, Zhuohao Yu, Mengyuan Sun, Zhibang Yang, Wei Wang, Hongrui Jia, Shikun Zhang, Wei Ye arxiv

Large language models are increasingly evaluated as interactive agents, yet standard agent benchmarks conflate two qualitatively distinct sources of success: semantic tool-use and interface-specific interaction pattern memorization. Because both mechanisms can yield identical task success on the original interface, benchmark scores alone are not identifiable evidence of environment-invariant capability. We propose PIPE, a protocol-level evaluation augmentation for diagnosing interface reliance by minimally rewriting environment interfaces while preserving task semantics and execution behavior. Across 16 environments from AgentBench and AgentGym and a range of open-source and API-based agents, PIPE reveals that trajectory-SFT substantially amplifies interface shortcutting: trained agents degrade sharply under minimal interface rewrites, while non-trajectory-trained models remain largely stable. We further introduce Interface Reliance (IR), a counterbalanced alias-based metric that quantifies preference for training-time interfaces, and show that interface shortcutting exhibits environment-dependent, non-monotonic training dynamics that remain invisible under standard evaluation. Our code is available at https://anonymous.4open.science/r/What-Do-Agents-Learn-from-Trajectory-SFT-Semantics-or-Interfaces--0831/.

📄 PDF Abstract BibTeX arXiv:2602.01611

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Tell me why: Training preferences-based RL with human preferences and step-level explanations

2024-05-23 · Jakob Karalus

Human-in-the-loop reinforcement learning allows the training of agents through various interfaces, even for non-expert humans. Recently, preference-based methods (PbRL), where the human has to give his preference over tw…

WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces

2026-06-08 · Wanli Li, Bowen Zhou, Yunyao Yu, Zhou Xu 외 arxiv

Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and external tools. Existing benchmarks, however, often evaluate these inte…

BGM: Building a Dynamic Guidance Map without Visual Images for Trajectory Prediction

2020-10-08 · Beihao Xia, Conghao Wong, Heng Li, Shiming Chen 외

Visual images usually contain the informative context of the environment, thereby helping to predict agents' behaviors. However, they hardly impose the dynamic effects on agents' actual behaviors due to the respectively …

DecoderTrajectory Prediction

Scalable Environments Drive Generalizable Agents

2026-05-18 · Jiayi Zhang, Fanqi Kong, Guibin Zhang, Maojia Song 외 arxiv

Generalizable agents should adapt to diverse tasks and unseen environments beyond their training distribution. This position paper argues that such generalization requires environment scaling: expanding the distribution …

Recent Advances in Text-to-SQL: A Survey of What We Have and What We Expect

2022-08-22 · COLING 2022 10 · Naihao Deng, Yulong Chen, Yue Zhang

Text-to-SQL has attracted attention from both the natural language processing and database communities because of its ability to convert the semantics in natural language into SQL queries and its practical application in…

SurveyText to SQLText-To-SQL