paper-with-me

홈 › Papers

PARL: Prompt-based Agents for Reinforcement Learning

2025-10-24 · Yarik Menchaca Resendiz, Roman Klinger arxiv

Large language models (LLMs) have demonstrated high performance on tasks expressed in natural language, particularly in zero- or few-shot settings. These are typically framed as supervised (e.g., classification) or unsupervised (e.g., clustering) problems. However, limited work evaluates LLMs as agents in reinforcement learning (RL) tasks (e.g., playing games), where learning occurs through interaction with an environment and a reward system. While prior work focused on representing tasks that rely on a language representation, we study structured, non-linguistic reasoning - such as interpreting positions in a grid world. We therefore introduce PARL (Prompt-based Agent for Reinforcement Learning), a method that uses LLMs as RL agents through prompting, without any fine-tuning. PARL encodes actions, states, and rewards in the prompt, enabling the model to learn through trial-and-error interaction. We evaluate PARL on three standard RL tasks that do not entirely rely on natural language. We show that it can match or outperform traditional RL agents in simple environments by leveraging pretrained knowledge. However, we identify performance limitations in tasks that require complex mathematical operations or decoding states and actions.

📄 PDF Abstract BibTeX arXiv:2510.21306

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

GraspARL: Dynamic Grasping via Adversarial Reinforcement Learning

2022-03-04 · Tianhao Wu, Fangwei Zhong, Yiran Geng, Hongchen Wang 외

Grasping moving objects, such as goods on a belt or living animals, is an important but challenging task in robotics. Conventional approaches rely on a set of manually defined object motion patterns for training, resulti…

Objectreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Open-World Dynamic Prompt and Continual Visual Representation Learning

2024-09-09 · Youngeun Kim, Jun Fang, Qin Zhang, Zhaowei Cai 외

The open world is inherently dynamic, characterized by ever-evolving concepts and distributions. Continual learning (CL) in this dynamic open-world environment presents a significant challenge in effectively generalizing…

Continual LearningImage RetrievalPrompt LearningRepresentation Learning

Predictable Reinforcement Learning Dynamics through Entropy Rate Minimization

2023-11-30 · Daniel Jarne Ornia, Giannis Delimpaltadakis, Jens Kober, Javier Alonso-Mora

In Reinforcement Learning (RL), agents have no incentive to exhibit predictable behaviors, and are often pushed (through e.g. policy entropy regularisation) to randomise their actions in favor of exploration. This often …

Policy Gradient Methodsreinforcement-learningReinforcement LearningReinforcement Learning (RL)

adaPARL: Adaptive Privacy-Aware Reinforcement Learning for Sequential-Decision Making Human-in-the-Loop Systems

2023-03-07 · Mojtaba Taherisadr, Stelios Andrew Stavroulakis, Salma Elmalaki

Reinforcement learning (RL) presents numerous benefits compared to rule-based approaches in various applications. Privacy concerns have grown with the widespread use of RL trained with privacy-sensitive data in IoT devic…

Decision MakingReinforcement Learning (RL)Sequential Decision Making

Preference-Aware Rubric Learning for Personalized Evaluation

2026-05-29 · Yilun Qiu, Xiaoyan Zhao, Yang Zhang, Yuxin Chen 외 arxiv

As Large Language Models (LLMs) evolve from general-purpose assistants to user-centric agents, personalization has become central to aligning model behavior with individual preferences, making the evaluation of personali…

Reinforcement LearningText Generation