paper-with-me

홈 › Papers

Emerging from Ground: Addressing Intent Deviation in Tool-Using Agents via Deriving Real Calls into Virtual Trajectories

2026-01-21 · Qian Xiong, Yuekai Huang, Bo Yang, Yujia Zheng, Tianhao Li, Ziyou Jiang, Zhiyuan Chang, Zhaoyang Li, Huanxiang Feng, Mingyang Li arxiv

LLMs have advanced tool-using agents for real-world applications, yet they often lead to unexpected behaviors or results. Beyond obvious failures, the subtle issue of "intent deviation" severely hinders reliable evaluation and performance improvement. Existing post-training methods generally leverage either real system samples or virtual data simulated by LLMs. However, the former is costly due to reliance on hand-crafted user requests, while the latter suffers from distribution shift from the real tools in the wild. Additionally, both methods lack negative samples tailored to intent deviation scenarios, hindering effective guidance on preference learning. We introduce RISE, a "Real-to-Virtual" method designed to mitigate intent deviation. Anchoring on verified tool primitives, RISE synthesizes virtual trajectories and generates diverse negative samples through mutation on critical parameters. With synthetic data, RISE fine-tunes backbone LLMs via the two-stage training for intent alignment. Evaluation results demonstrate that data synthesized by RISE achieve promising results in eight metrics covering user requires, execution trajectories and agent responses. Integrating with training, RISE achieves an average 35.28% improvement in Acctask (task completion) and 23.27% in Accintent (intent alignment), outperforming SOTA baselines by 1.20--42.09% and 1.17--54.93% respectively.

📄 PDF Abstract BibTeX arXiv:2601.15120

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ASTRA-bench: Evaluating Tool-Use Agent Reasoning and Action Planning with Personal User Context

2026-03-02 · Zidi Xiu, David Q. Sun, Kevin Cheng, Maitrik Patel 외 arxiv

Next-generation AI must manage vast personal data, diverse tools, and multi-step reasoning, yet most benchmarks remain context-free and single-turn. We present ASTRA-bench (Assistant Skills in Tool-use, Reasoning \& Acti…

Physics-Grounded Multi-Agent Architecture for Traceable, Risk-Aware Human-AI Decision Support in Manufacturing

2026-05-05 · Danny Hoang, Ryan Matthiessen, Christopher Miller, Nasir Mannan 외 arxiv

High-precision CNC machining of free-form aerospace components requires bounded compensations informed by inspection, simulation, and process knowledge. Off-the-shelf large language model (LLM) assistants can generate te…

Emerging Challenges of Integrating Solar PV in the Ireland and Northern Ireland Power Systems

2024-04-06 · Taulant Kerci, Manuel Hurtado, Simon Tweed, Marta Val Escudero 외

This paper discusses emerging operational challenges associated with the integration of solar photovoltaic (PV) in the All-Island power system (AIPS) of Ireland and Northern Ireland. These include the impact of solar PV …

Projecting the Emerging Mindset of SWE Agent by Launching a Wild Code Understanding Journey

2026-06-07 · Zhengyi Zhuo, Yan Liu arxiv

Software engineering agents (SWE agents) increasingly work through tool-mediated trajectories in real repositories, yet their behavior remains difficult to characterize in concrete, observable terms. These trajectories r…

Mind-Brush: Integrating Agentic Cognitive Search and Reasoning into Image Generation

2026-02-02 · Jun He, Junyan Ye, Zilong Huang, Dongzhi Jiang 외 arxiv

While text-to-image generation has achieved unprecedented fidelity, the vast majority of existing models function fundamentally as static text-to-pixel decoders. Consequently, they often fail to grasp implicit user inten…

Text-to-Image Generation