paper-with-me

홈 › Papers

EPO: Hierarchical LLM Agents with Environment Preference Optimization

2024-08-28 · Qi Zhao, Haotian Fu, Chen Sun, George Konidaris

Long-horizon decision-making tasks present significant challenges for LLM-based agents due to the need for extensive planning over multiple steps. In this paper, we propose a hierarchical framework that decomposes complex tasks into manageable subgoals, utilizing separate LLMs for subgoal prediction and low-level action generation. To address the challenge of creating training signals for unannotated datasets, we develop a reward model that leverages multimodal environment feedback to automatically generate reward signals. We introduce Environment Preference Optimization (EPO), a novel method that generates preference signals from the environment's feedback and uses them to train LLM-based agents. Extensive experiments on ALFRED demonstrate the state-of-the-art performance of our framework, achieving first place on the ALFRED public leaderboard and showcasing its potential to improve long-horizon decision-making in diverse environments.

📄 PDF Abstract BibTeX arXiv:2408.16090

Code (1)

kevinz8866/epo 공식 구현 pytorch

Tasks

Action GenerationDecision Making

Similar Papers 제목 키워드 기반

Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM Agents

2025-09-26 · Heyang Gao, Zexu Sun, Erxue Min, Hengyi Cai 외 arxiv

Large Language Models (LLMs) as autonomous agents are increasingly tasked with solving complex, long-horizon problems. Aligning these agents via preference-based offline methods like Direct Preference Optimization (DPO) …

Reasoning in a Hierarchical System with Missing Group Size Information

2018-02-07 · Subhash Kak

The paper analyzes the problem of judgments or preferences subsequent to initial analysis by autonomous agents in a hierarchical system where the higher level agents does not have access to group size information. We pro…

LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents

2026-05-28 · Xiaoxuan Peng, Kaiqi Zhang, Xinyu Lu, Boxi Cao 외 arxiv

Mastering terminal environments requires language agents capable of multi-step planning, feedback-grounded execution, and dynamic state adaptation. However, training such agents is currently bottlenecked by a reliance on…

DIPPER: Direct Preference Optimization to Accelerate Primitive-Enabled Hierarchical Reinforcement Learning

2024-06-16 · Utsav Singh, Souradip Chakraborty, Wesley A. Suttle, Brian M. Sadler 외

Learning control policies to perform complex robotics tasks from human preference data presents significant challenges. On the one hand, the complexity of such tasks typically requires learning policies to perform a vari…

Computational EfficiencyHierarchical Reinforcement Learningreinforcement-learningReinforcement Learning

Coordination Requires Simplification: Thermodynamic Bounds on Multi-Objective Compromise in Natural and Artificial Intelligence

2025-09-27 · Atma Anand arxiv

Information-processing systems that coordinate multiple agents and objectives face fundamental thermodynamic constraints. We show that solutions with maximum utility to act as coordination focal points have a much higher…

Reinforcement Learning