paper-with-me

홈 › Papers

Towards a Design Guideline for RPA Evaluation: A Survey of Large Language Model-Based Role-Playing Agents

2025-02-18 · Chaoran Chen, Bingsheng Yao, Ruishi Zou, Wenyue Hua, Weimin Lyu, Yanfang Ye, Toby Jia-Jun Li, Dakuo Wang

Role-Playing Agent (RPA) is an increasingly popular type of LLM Agent that simulates human-like behaviors in a variety of tasks. However, evaluating RPAs is challenging due to diverse task requirements and agent designs. This paper proposes an evidence-based, actionable, and generalizable evaluation design guideline for LLM-based RPA by systematically reviewing 1,676 papers published between Jan. 2021 and Dec. 2024. Our analysis identifies six agent attributes, seven task attributes, and seven evaluation metrics from existing literature. Based on these findings, we present an RPA evaluation design guideline to help researchers develop more systematic and consistent evaluation methods.

📄 PDF Abstract BibTeX arXiv:2502.13012

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Survey-aware Machine Learning: A Guideline for Valid Population Health Inference based on Scoping Review

2026-05-09 · YongKyung Oh, Henry W. Zheng, Jeffrey Feng, Alex A. T. Bui arxiv

Machine Learning (ML) models trained on complex health surveys such as the National Health and Nutrition Examination Survey (NHANES) often ignore primary sampling units, stratification variables, and sampling weights. Th…

Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey

2025-05-21 · Chih-Kai Yang, Neo S. Ho, Hung-Yi Lee

With advancements in large audio-language models (LALMs), which enhance large language models (LLMs) with auditory capabilities, these models are expected to demonstrate universal proficiency across various auditory task…

FairnessSurvey

On Context Span Needed for Machine Translation Evaluation

2020-05-01 · LREC 2020 5 · Sheila Castilho, Maja Popovi{\'c}, Andy Way

Despite increasing efforts to improve evaluation of machine translation (MT) by going beyond the sentence level to the document level, the definition of what exactly constitutes a {``}document level{''} is still not clea…

Machine TranslationSentenceTranslation

Human Evaluation of Creative NLG Systems: An Interdisciplinary Survey on Recent Papers

2021-07-31 · ACL (GEM) 2021 8 · Mika Hämäläinen, Khalid Alnajjar

We survey human evaluation in papers presenting work on creative natural language generation that have been published in INLG 2020 and ICCC 2020. The most typical human evaluation method is a scaled survey, typically on …

SurveyText Generation

AI-Driven Test Case Generation from Natural Language Requirements: A Survey of Techniques and Research Gaps

2026-06-04 · Orimoloye Folorunsho, Hassan Reza arxiv

Software testing is critical for verifying that systems meet specified requirements, yet remains among the most time-consuming and expensive activities in development. Requirements-based test generation allows test cases…