paper-with-me

홈 › Papers

Hierarchical Reward Design from Language: Enhancing Alignment of Agent Behavior with Human Specifications

2026-02-20 · Zhiqin Qian, Ryan Diaz, Sangwon Seo, Vaibhav Unhelkar arxiv

When training artificial intelligence (AI) to perform tasks, humans often care not only about whether a task is completed but also how it is performed. As AI agents tackle increasingly complex tasks, aligning their behavior with human-provided specifications becomes critical for responsible AI deployment. Reward design provides a direct channel for such alignment by translating human expectations into reward functions that guide reinforcement learning (RL). However, existing methods are often too limited to capture nuanced human preferences that arise in long-horizon tasks. Hence, we introduce Hierarchical Reward Design from Language (HRDL): a problem formulation that extends classical reward design to encode richer behavioral specifications for hierarchical RL agents. We further propose Language to Hierarchical Rewards (L2HR) as a solution to HRDL. Experiments show that AI agents trained with rewards designed via L2HR not only complete tasks effectively but also better adhere to human specifications. Together, HRDL and L2HR advance the research on human-aligned AI agents.

📄 PDF Abstract BibTeX arXiv:2602.18582

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

ALaRM: Align Language Models via Hierarchical Rewards Modeling

2024-03-11 · Yuhang Lai, Siyuan Wang, Shujun Liu, Xuanjing Huang 외

We introduce ALaRM, the first framework modeling hierarchical rewards in reinforcement learning from human feedback (RLHF), which is designed to enhance the alignment of large language models (LLMs) with human preference…

Long Form Question AnsweringMachine TranslationQuestion AnsweringText Generation

Language Chain in Alignment: Cross-lingual Ranking Preference Optimization

2026-08-24 · Seungyoon Lee, Minhyuk Kim, Jungseob Lee, Heuiseok Lim arxiv

The alignment of Large Language Models heavily relies on English-centric high-quality preference data, which often leads to suboptimal performance in other languages. In this paper, we propose Cross-lingual Ranking Prefe…

HAF-RM: A Hybrid Alignment Framework for Reward Model Training

2024-07-04 · Shujun Liu, Xiaoyu Shen, Yuhang Lai, Siyuan Wang 외

The reward model has become increasingly important in alignment, assessment, and data construction for large language models (LLMs). Most existing researchers focus on enhancing reward models through data improvements, f…

Aligning Anime Video Generation with Human Feedback

2025-04-14 · Bingwen Zhu, Yudong Jiang, Baohan Xu, Siqian Yang 외

Anime video generation faces significant challenges due to the scarcity of anime data and unusual motion patterns, leading to issues such as motion distortion and flickering artifacts, which result in misalignment with h…

Video Generation

TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards

2026-05-19 · Mingxuan Cui, Jingpu Yang, Fengxian Ji, Qian Jiang 외 arxiv

Faithful text rendering remains a persistent weakness of large text-to-image generative models, as it requires both semantic instruction following and fine-grained glyph-level structure. Prior methods often improve this …

Instruction Following