paper-with-me

홈 › Papers

AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

2024-06-06 · Zhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong, Honglin Guo, Junzhe Wang, Dingwen Yang, Chenyang Liao, Xin Guo, wei he, Songyang Gao, Lu Chen, Rui Zheng, Yicheng Zou, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjing Huang, Zuxuan Wu, Yu-Gang Jiang

Building generalist agents that can handle diverse tasks and evolve themselves across different environments is a long-term goal in the AI community. Large language models (LLMs) are considered a promising foundation to build such agents due to their generalized capabilities. Current approaches either have LLM-based agents imitate expert-provided trajectories step-by-step, requiring human supervision, which is hard to scale and limits environmental exploration; or they let agents explore and learn in isolated environments, resulting in specialist agents with limited generalization. In this paper, we take the first step towards building generally-capable LLM-based agents with self-evolution ability. We identify a trinity of ingredients: 1) diverse environments for agent exploration and learning, 2) a trajectory set to equip agents with basic capabilities and prior knowledge, and 3) an effective and scalable evolution method. We propose AgentGym, a new framework featuring a variety of environments and tasks for broad, real-time, uni-format, and concurrent agent exploration. AgentGym also includes a database with expanded instructions, a benchmark suite, and high-quality trajectories across environments. Next, we propose a novel method, AgentEvol, to investigate the potential of agent self-evolution beyond previously seen data across tasks and environments. Experimental results show that the evolved agents can achieve results comparable to SOTA models. We release the AgentGym suite, including the platform, dataset, benchmark, checkpoints, and algorithm implementations. The AgentGym suite is available on https://github.com/WooooDyy/AgentGym.

📄 PDF Abstract BibTeX arXiv:2406.04151

Code (1)

woooodyy/agentgym 공식 구현

Tasks

Language ModelingLanguage ModellingLarge Language Model

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

2025-09-10 · Zhiheng Xi, Jixuan Huang, Chenyang Liao, Baodai Huang 외 arxiv

Developing autonomous LLM agents capable of making a series of intelligent decisions to solve complex, real-world tasks is a fast-evolving frontier. Like human cognitive development, agents are expected to acquire knowle…

Reinforcement LearningDecision Making

MedAgentGym: Training LLM Agents for Code-Based Medical Reasoning at Scale

2025-06-04 · ran Xu, Yuchen Zhuang, Yishan Zhong, Yue Yu 외

We introduce MedAgentGYM, the first publicly available training environment designed to enhance coding-based medical reasoning capabilities in large language model (LLM) agents. MedAgentGYM comprises 72,413 task instance…

BenchmarkingLanguage ModelingLanguage ModellingLarge Language Model+1

AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments

2026-07-06 · Zhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang 외 arxiv

Language agents, i.e., LLM agents, progress rapidly and are increasingly deployed in production environments. This trend underscores the urgent need for rigorous and realistic evaluations. However, most existing benchmar…

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents

2026-02-13 · Yujiong Shen, Yajie Yang, Zhiheng Xi, Binze Hu 외 arxiv

Scientific reasoning inherently demands integrating sophisticated toolkits to navigate domain-specific knowledge. Yet, current benchmarks largely overlook agents' ability to orchestrate tools for such rigorous workflows.…

What Do Agents Learn from Trajectory-SFT: Semantics or Interfaces?

2026-02-02 · Weizheng Gu, Chengze Li, Zhuohao Yu, Mengyuan Sun 외 arxiv

Large language models are increasingly evaluated as interactive agents, yet standard agent benchmarks conflate two qualitatively distinct sources of success: semantic tool-use and interface-specific interaction pattern m…