paper-with-me

Papers

Skill Check: Some Considerations on the Evaluation of Gamemastering Models for Role-playing Games

2023-09-24 · Santiago Góngora, Luis Chiruzzo, Gonzalo Méndez, Pablo Gervás

In role-playing games a Game Master (GM) is the player in charge of the game, who must design the challenges the players face and narrate the outcomes of their actions. In this work we discuss some challenges to model GMs from an Interactive Storytelling and Natural Language Processing perspective. Following those challenges we propose three test categories to evaluate such dialogue systems, and we use them to test ChatGPT, Bard and OpenAssistant as out-of-the-box GMs.

📄 PDF Abstract BibTeX arXiv:2309.13702

Code (1)

sgongora27/skill-check-gm-tests 공식 구현

Similar Papers 제목 키워드 기반

Data science skills for referees: I Biological X-ray crystallography

2017-04-28

Since there is now a growing wish by referees to judge the underpinning data for a submitted article it is timely to provide a summary of the data evaluation checks required to be done by a referee. As these checks will …

Articles

SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use

2026-07-02 · Jiayin Zhu, Kelong Mao, Yudong Guo, Dengbo He 외 arxiv

Skills are becoming a reusable operational layer for LLM agents, encoding SOPs, domain rules, tool workflows, scripts, and validation routines. In realistic skill repositories, overlapping skills make reliable skill-use …

SkillAudit: From Fixed-Suite Benchmarking to Skill-Centered Assessment

2026-06-21 · Dexu Yu, Youhua Li, Zhaoyang Guan, Xianhao Lin 외 arxiv

Agent skills have become a practical way to extend large language model agents, but the growing skill ecosystem still lacks a reliable way to judge whether a skill is worth deploying. Existing evaluation methods remain l…

Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill

2026-06-02 · Tao Chen, Gangwei Jiang, Pengyu Cheng, Siyuan Huang 외 arxiv

Reward models (RMs) provide critical feedback signals for LLM post-training, notably in reinforced fine-tuning (RFT) and reinforcement learning (RL) pipelines. However, current reward evaluation relies on heterogeneous c…

Reinforcement Learning

Trusted Source Alignment in Large Language Models

2023-11-12 · Vasilisa Bashlovkina, Zhaobin Kuang, Riley Matthews, Edward Clifford 외

Large language models (LLMs) are trained on web-scale corpora that inevitably include contradictory factual information from sources of varying reliability. In this paper, we propose measuring an LLM property called trus…

ArticlesFact Checking