paper-with-me

홈 › Papers

PingPong: A Benchmark for Role-Playing Language Models with User Emulation and Multi-Model Evaluation

2024-09-10 · Ilya Gusev

We introduce a benchmark for evaluating the role-playing capabilities of language models. Our approach leverages language models themselves to emulate users in dynamic, multi-turn conversations and to assess the resulting dialogues. The framework consists of three main components: a player model that assumes a specific character role, an interrogator model that simulates user behavior, and several judge models that evaluate conversation quality. We conducted experiments comparing automated evaluations with human annotations to validate our approach, demonstrating strong correlations across multiple criteria. This work provides a foundation for a robust and dynamic evaluation of the model capabilities in interactive scenarios.

📄 PDF Abstract BibTeX arXiv:2409.06820

Code (1)

ilyagusev/ping_pong_bench 공식 구현

Similar Papers 제목 키워드 기반

RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models

2023-10-01 · Zekun Moore Wang, Zhongyuan Peng, Haoran Que, Jiaheng Liu 외

The advent of Large Language Models (LLMs) has paved the way for complex tasks such as role-playing, which enhances user interactions by enabling models to imitate various characters. However, the closed-source nature of…

Benchmarking

PingPong: A Natural Benchmark for Multi-Turn Code-Switching Dialogues

2026-01-24 · Mohammad Rifqi Farhansyah, Hanif Muhammad Zhafran, Farid Adilazuarda, Shamsuddeen Hassan Muhammad 외 arxiv

Code-switching is a widespread practice among the world's multilingual majority, yet few benchmarks accurately reflect its complexity in everyday communication. We present PingPong, a benchmark for natural multi-party co…

Question Answering

RMTBench: Benchmarking LLMs Through Multi-Turn User-Centric Role-Playing

2025-07-27 · Hao Xiang, Tianyi Tang, Yang Su, Bowen Yu 외 arxiv

Recent advancements in Large Language Models (LLMs) have shown outstanding potential for role-playing applications. Evaluating these capabilities is becoming crucial yet remains challenging. Existing benchmarks mostly ad…

TimeChara: Evaluating Point-in-Time Character Hallucination of Role-Playing Large Language Models

2024-05-28 · Jaewoo Ahn, Taehyun Lee, Junyoung Lim, Jin-Hwa Kim 외

While Large Language Models (LLMs) can serve as agents to simulate human behaviors (i.e., role-playing agents), we emphasize the importance of point-in-time role-playing. This situates characters at specific moments in t…

Hallucination

RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue

2026-09-15 · Yuqi Wang, Fengyuan Liu, Haochen Luo, Zhiqi Yu 외 arxiv

Speech-to-speech dialogue models increasingly support persona control, yet existing spoken role-playing benchmarks remain largely character-centric and short-horizon. This leaves open whether spoken dialogue models can s…