paper-with-me

홈 › Papers

Playpen: An Environment for Exploring Learning Through Conversational Interaction

2025-04-11 · Nicola Horst, Davide Mazzaccara, Antonia Schmidt, Michael Sullivan, Filippo Momentè, Luca Franceschetti, Philipp Sadler, Sherzod Hakimov, Alberto Testoni, Raffaella Bernardi, Raquel Fernández, Alexander Koller, Oliver Lemon, David Schlangen, Mario Giulianelli, Alessandro Suglia

Are we running out of learning signal? Predicting the next word in an existing text has turned out to be a powerful signal, at least at scale. But there are signs that we are running out of this resource. In recent months, interaction between learner and feedback-giver has come into focus, both for "alignment" (with a reward model judging the quality of instruction following attempts) and for improving "reasoning" (process- and outcome-based verifiers judging reasoning steps). In this paper, we explore to what extent synthetic interaction in what we call Dialogue Games -- goal-directed and rule-governed activities driven predominantly by verbal actions -- can provide a learning signal, and how this signal can be used. We introduce an environment for producing such interaction data (with the help of a Large Language Model as counterpart to the learner model), both offline and online. We investigate the effects of supervised fine-tuning on this data, as well as reinforcement learning setups such as DPO, and GRPO; showing that all of these approaches achieve some improvements in in-domain games, but only GRPO demonstrates the ability to generalise to out-of-domain games as well as retain competitive performance in reference-based tasks. We release the framework and the baseline training setups in the hope that this can foster research in this promising new direction.

📄 PDF Abstract BibTeX arXiv:2504.08590

Code (1)

lm-playpen/playpen 공식 구현

Tasks

Instruction FollowingLarge Language Model

Methods 이 논문이 사용한 방법론

DPO 설명 없음

Similar Papers 제목 키워드 기반

Exploring Interaction Patterns for Debugging: Enhancing Conversational Capabilities of AI-assistants

2024-02-09 · Bhavya Chopra, Yasharth Bajpai, Param Biyani, Gustavo Soares 외

The widespread availability of Large Language Models (LLMs) within Integrated Development Environments (IDEs) has led to their speedy adoption. Conversational interactions with LLMs enable programmers to obtain natural l…

Fault localization

ConvoKit: A Toolkit for the Analysis of Conversations

2020-05-08 · SIGDIAL (ACL) 2020 7 · Jonathan P. Chang, Caleb Chiam, Liye Fu, Andrew Z. Wang 외

This paper describes the design and functionality of ConvoKit, an open-source toolkit for analyzing conversations and the social interactions embedded within. ConvoKit provides an unified framework for representing and m…

"I Like Sunnie More Than I Expected!": Exploring User Expectation and Perception of an Anthropomorphic LLM-based Conversational Agent for Well-Being Support

2024-05-22 · Siyi Wu, Julie Y. A. Cachia, Feixue Han, Bingsheng Yao 외

The human-computer interaction (HCI) research community has a longstanding interest in exploring the mismatch between users' actual experiences and expectation toward new technologies, for instance, large language models…

Recommendation Systems

Voice Interaction With Conversational AI Could Facilitate Thoughtful Reflection and Substantive Revision in Writing

2025-04-11 · Jiho Kim, Philippe Laban, Xiang 'Anthony' Chen, Kenneth C. Arnold

Writing well requires not only expressing ideas but also refining them through revision, a process facilitated by reflection. Prior research suggests that feedback delivered through dialogues, such as those in writing ce…

What Do You Mean? Exploring How Humans and AI Interact with Symbols and Meanings in Their Interactions

2025-10-06 · Reza Habibi, Seung Wan Ha, Zhiyu Lin, Atieh Kashani 외 arxiv

Meaningful human-AI collaboration requires more than processing language; it demands a deeper understanding of symbols and their socially constructed meanings. While humans naturally interpret symbols through social inte…