paper-with-me

홈 › Papers

The End of Reward Engineering: How LLMs Are Redefining Multi-Agent Coordination

2026-01-13 · Haoran Su, Yandong Sun, Congjia Yu arxiv

Reward engineering, the manual specification of reward functions to induce desired agent behavior, remains a fundamental challenge in multi-agent reinforcement learning. This difficulty is amplified by credit assignment ambiguity, environmental non-stationarity, and the combinatorial growth of interaction complexity. We argue that recent advances in large language models (LLMs) point toward a shift from hand-crafted numerical rewards to language-based objective specifications. Prior work has shown that LLMs can synthesize reward functions directly from natural language descriptions (e.g., EUREKA) and adapt reward formulations online with minimal human intervention (e.g., CARD). In parallel, the emerging paradigm of Reinforcement Learning from Verifiable Rewards (RLVR) provides empirical evidence that language-mediated supervision can serve as a viable alternative to traditional reward engineering. We conceptualize this transition along three dimensions: semantic reward specification, dynamic reward adaptation, and improved alignment with human intent, while noting open challenges related to computational overhead, robustness to hallucination, and scalability to large multi-agent systems. We conclude by outlining a research direction in which coordination arises from shared semantic representations rather than explicitly engineered numerical signals.

📄 PDF Abstract BibTeX arXiv:2601.08237

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-agent Reinforcement Learning

Similar Papers 제목 키워드 기반

GOV-REK: Governed Reward Engineering Kernels for Designing Robust Multi-Agent Reinforcement Learning Systems

2024-04-01 · Ashish Rana, Michael Oesterle, Jannik Brinkmann

For multi-agent reinforcement learning systems (MARLS), the problem formulation generally involves investing massive reward engineering effort specific to a given problem. However, this effort often cannot be translated …

Multi-agent Reinforcement Learning

LLMs: A Game-Changer for Software Engineers?

2024-11-01 · Md Asraful Haque

Large Language Models (LLMs) like GPT-3 and GPT-4 have emerged as groundbreaking innovations with capabilities that extend far beyond traditional AI applications. These sophisticated models, trained on massive datasets, …

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

2025-06-13 · Jeff Da, Clinton Wang, Xiang Deng, Yuntao Ma 외

Reinforcement Learning from Verifiable Rewards (RLVR) has been widely adopted as the de facto method for enhancing the reasoning capabilities of large language models and has demonstrated notable success in verifiable do…

MathNavigate

Act-Observe-Rewrite: Multimodal Coding Agents as In-Context Policy Learners for Robot Manipulation

2026-03-03 · Vaishak Kumar arxiv

Can a multimodal language model learn to manipulate physical objects by reasoning about its own failures-without gradient updates, demonstrations, or reward engineering? We argue the answer is yes, under conditions we ch…

Robot ManipulationCode Generation

PCGRLLM: Large Language Model-Driven Reward Design for Procedural Content Generation Reinforcement Learning

2025-02-15 · In-Chang Baek, Sung-Hyun Kim, Sam Earle, Zehua Jiang 외

Reward design plays a pivotal role in the training of game AIs, requiring substantial domain-specific knowledge and human effort. In recent years, several studies have explored reward generation for training game agents …

Language ModelingLanguage ModellingLarge Language ModelPrompt Engineering