paper-with-me

홈 › Papers

Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning

2026-07-26 · Wenxuan Zhang, Yuhui Wang, Donggang Jia, Xiaoqian Shen, Jian Ding, Ivan Viola, Jürgen Schmidhuber, Mohamed Elhoseiny arxiv

Large Vision-Language Models (VLMs) now act as agents in interactive environments, where success requires coherent reasoning and decision-making across turns. Although end-to-end training in agentic environments can improve such multi-turn decision-making abilities, current methods mainly rely on either token-wise optimization over concatenated token trajectories or turn-wise optimization with uniform within-turn credit. In this work, we establish theoretical formulations for the two levels of optimization and derive a hybrid advantage that serves both objectives. Furthermore, with an appropriate choice of discount factor and learning target, we prove that a unified critic model can estimate values for both turn-wise and token-wise. As such, we propose HyGAE, an actor-critic framework that jointly optimizes token- and turn-level objectives with the hybrid advantage and unified critic. We conduct extensive evaluations of HyGAE across five multi-turn decision-making environments, where it achieves an average success rate of 91% and a significant improvement of 10% over other methods. Furthermore, we provide an in-depth analysis showing that the exact analytic form of the hybrid advantage and return is crucial for optimization. Project Page: https://wx-zhang.github.io/hygae-web/.

📄 PDF Abstract BibTeX arXiv:2607.23605

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Enhancing Steering Estimation with Semantic-Aware GNNs

2025-03-21 · Fouad Makiyeh, Huy-Dung Nguyen, Patrick Chareyre, Ramin Hasani 외

Steering estimation is a critical task in autonomous driving, traditionally relying on 2D image-based models. In this work, we explore the advantages of incorporating 3D spatial information through hybrid architectures t…

Autonomous Drivinggraph constructionGraph Neural Network

BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning

2026-06-01 · Shannon Serrao, Soumitra Chatterjee, Dorina Strori, Abhishek Sharma 외 arxiv

Enterprise AI systems that translate natural language into SQL queries and orchestrate multi-step agentic reasoning pipelines require evaluation approaches fundamentally different from academic benchmarks. Spider and BIR…

Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning

2026-06-22 · Yunan Wang, Minghui Song, Zihan Zhang, Shaohan Huang 외 arxiv

Group-based Reinforcement Learning (RL) has significantly enhanced Large Language Models (LLMs) in agentic scenarios. To achieve finer-grained policy updates, recent agentic RL frameworks have shifted from trajectory-lev…

Reinforcement Learning

AutoForge: Automated Environment Synthesis for Agentic Reinforcement Learning

2025-12-28 · Shihao Cai, Runnan Fang, Jialong Wu, Baixuan Li 외 arxiv

Conducting reinforcement learning (RL) in simulated environments offers a cost-effective and highly scalable way to enhance language-based agents. However, previous work has been limited to semi-automated environment syn…

Reinforcement LearningDomain Generalization

Transforming the Hybrid Cloud for Emerging AI Workloads

2024-11-20 · Deming Chen, Alaa Youssef, Ruchi Pendse, André Schleife 외

This white paper, developed through close collaboration between IBM Research and UIUC researchers within the IIDAI Institute, envisions transforming hybrid cloud systems to meet the growing complexity of AI workloads thr…

Model Optimizationscientific discoveryWeather Forecasting