paper-with-me

홈 › Papers

Beyond Static Evaluation: Building Simulation Environments for Scalable Agentic Reinforcement Learning

2026-07-07 · Akshay Arora, Ishan Nigam, Ashutosh Aggarwal, Shefali Bansal, Krishna Singh, Sweta Kumari, Nikhil Mittal, Shariq Farhan, Siddarth Malreddy arxiv

As Large Language Models (LLMs) evolve into autonomous agents, traditional static evaluation fails to capture multi-step decision-making. We introduce AgenticAI-Supervisor, an API and UI-driven RL Gym environment that decouples environment creation from scalable execution. By moving to verifiable execution outcomes, the platform generates high-fidelity traces and applies multi-dimensional reward shaping. Critically, our framework mitigates reward hacking through rigorous internal state validation and testing. This work provides a first look at our platform's core capabilities through a Customer Support Agent case study demonstrating a consistent closed-loop feedback for model optimization. Future work will focus on advanced features such as Computer Use, Tool Use, automated "stumping", and edge-case generation.

📄 PDF Abstract BibTeX arXiv:2607.05773

Code (2)

Aaron617/agent-arXiv-daily ★ 10
Tavish9/awesome-daily-AI-arxiv ★ 111

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

IndoorR2X: Indoor Robot-to-Everything Coordination with LLM-Driven Planning

2026-03-20 · Fan Yang, Soumya Teotia, Shaunak A. Mehta, Prajit KrisshnaKumar 외 arxiv

Although robot-to-robot (R2R) communication improves indoor scene understanding beyond what a single robot can achieve, R2R alone cannot overcome partial observability without substantial exploration overhead or scaling …

Robot Task PlanningScene Understanding

Static Sandboxes Are Inadequate: Modeling Societal Complexity Requires Open-Ended Co-Evolution in LLM-Based Multi-Agent Simulations

2025-10-15 · Jinkun Chen, Sher Badshah, Xuemin Yu, Sijia Han arxiv

What if artificial agents could not just communicate, but also evolve, adapt, and reshape their worlds in ways we cannot fully predict? With llm now powering multi-agent systems and social simulations, we are witnessing …

DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation

2024-11-18 · CVPR 2025 1 · Tianyi Yan, Dongming Wu, Wencheng Han, Junpeng Jiang 외

Autonomous driving evaluation requires simulation environments that closely replicate actual road conditions, including real-world sensory data and responsive feedback loops. However, many existing simulations need to pr…

Autonomous DrivingDecision Making

B2RL: An open-source Dataset for Building Batch Reinforcement Learning

2022-09-30 · Hsin-Yu Liu, Xiaohan Fu, Bharathan Balaji, Rajesh Gupta 외

Batch reinforcement learning (BRL) is an emerging research area in the RL community. It learns exclusively from static datasets (i.e. replay buffers) without interaction with the environment. In the offline settings, exi…

Managementreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Efficient 3D Reconstruction, Streaming and Visualization of Static and Dynamic Scene Parts for Multi-client Live-telepresence in Large-scale Environments

2022-11-25 · Leif Van Holland, Patrick Stotko, Stefan Krumpen, Reinhard Klein 외

Despite the impressive progress of telepresence systems for room-scale scenes with static and dynamic scene entities, expanding their capabilities to scenarios with larger dynamic environments beyond a fixed size of a fe…

3D Reconstruction