paper-with-me

홈 › Papers

CharacterBench: Benchmarking Character Customization of Large Language Models

2024-12-16 · Jinfeng Zhou, Yongkang Huang, Bosi Wen, Guanqun Bi, Yuxuan Chen, Pei Ke, Zhuang Chen, Xiyao Xiao, Libiao Peng, Kuntian Tang, Rongsheng Zhang, Le Zhang, Tangjie Lv, Zhipeng Hu, Hongning Wang, Minlie Huang

Character-based dialogue (aka role-playing) enables users to freely customize characters for interaction, which often relies on LLMs, raising the need to evaluate LLMs' character customization capability. However, existing benchmarks fail to ensure a robust evaluation as they often only involve a single character category or evaluate limited dimensions. Moreover, the sparsity of character features in responses makes feature-focused generative evaluation both ineffective and inefficient. To address these issues, we propose CharacterBench, the largest bilingual generative benchmark, with 22,859 human-annotated samples covering 3,956 characters from 25 detailed character categories. We define 11 dimensions of 6 aspects, classified as sparse and dense dimensions based on whether character features evaluated by specific dimensions manifest in each response. We enable effective and efficient evaluation by crafting tailored queries for each dimension to induce characters' responses related to specific dimensions. Further, we develop CharacterJudge model for cost-effective and stable evaluations. Experiments show its superiority over SOTA automatic judges (e.g., GPT-4) and our benchmark's potential to optimize LLMs' character customization. Our repository is at https://github.com/thu-coai/CharacterBench.

📄 PDF Abstract BibTeX arXiv:2412.11912

Code (1)

thu-coai/characterbench 공식 구현

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

CharacterShot: Controllable and Consistent 4D Character Animation

2025-08-10 · Junyao Gao, Jiaxing Li, Wenran Liu, Yanhong Zeng 외 arxiv

In this paper, we propose \textbf{CharacterShot}, a controllable and consistent 4D character animation framework that enables any individual designer to create dynamic 3D characters (i.e., 4D character animation) from a …

RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models

2023-10-01 · Zekun Moore Wang, Zhongyuan Peng, Haoran Que, Jiaheng Liu 외

The advent of Large Language Models (LLMs) has paved the way for complex tasks such as role-playing, which enhances user interactions by enabling models to imitate various characters. However, the closed-source nature of…

Benchmarking

CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimization

2025-08-12 · Xinge Ye, Rui Wang, Yuchuan Wu, Victor Ma 외 arxiv

Reinforcement Learning Fine-Tuning (RLFT) has achieved notable success in tasks with objectively verifiable answers (e.g., code generation, mathematical reasoning), yet struggles with open-ended subjective tasks like rol…

Reinforcement LearningMathematical ReasoningCode Generation

LAVIS: A Library for Language-Vision Intelligence

2022-09-15 · Dongxu Li, Junnan Li, Hung Le, Guangsen Wang 외

We introduce LAVIS, an open-source deep learning library for LAnguage-VISion research and applications. LAVIS aims to serve as a one-stop comprehensive library that brings recent advancements in the language-vision field…

BenchmarkingImage CaptioningImage RetrievalMultimodal Deep Learning+7

Empirical Guidelines for Deploying LLMs onto Resource-constrained Edge Devices

2024-06-06 · Ruiyang Qin, Dancheng Liu, Chenhui Xu, Zheyu Yan 외

The scaling laws have become the de facto guidelines for designing large language models (LLMs), but they were studied under the assumption of unlimited computing resources for both training and inference. As LLMs are in…

BenchmarkingRAG