paper-with-me

홈 › Papers

Towards Personalized Evaluation of Large Language Models with An Anonymous Crowd-Sourcing Platform

2024-03-13 · Mingyue Cheng, Hao Zhang, Jiqian Yang, Qi Liu, Li Li, Xin Huang, Liwei Song, Zhi Li, Zhenya Huang, Enhong Chen

Large language model evaluation plays a pivotal role in the enhancement of its capacity. Previously, numerous methods for evaluating large language models have been proposed in this area. Despite their effectiveness, these existing works mainly focus on assessing objective questions, overlooking the capability to evaluate subjective questions which is extremely common for large language models. Additionally, these methods predominantly utilize centralized datasets for evaluation, with question banks concentrated within the evaluation platforms themselves. Moreover, the evaluation processes employed by these platforms often overlook personalized factors, neglecting to consider the individual characteristics of both the evaluators and the models being evaluated. To address these limitations, we propose a novel anonymous crowd-sourcing evaluation platform, BingJian, for large language models that employs a competitive scoring mechanism where users participate in ranking models based on their performance. This platform stands out not only for its support of centralized evaluations to assess the general capabilities of models but also for offering an open evaluation gateway. Through this gateway, users have the opportunity to submit their questions, testing the models on a personalized and potentially broader range of capabilities. Furthermore, our platform introduces personalized evaluation scenarios, leveraging various forms of human-computer interaction to assess large language models in a manner that accounts for individual user preferences and contexts. The demonstration of BingJian can be accessed at https://github.com/Mingyue-Cheng/Bingjian.

📄 PDF Abstract BibTeX arXiv:2403.08305

Code (1)

mingyue-cheng/bingjian 공식 구현

Tasks

Language Model EvaluationLanguage ModellingLarge Language Model

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Personalized Language Model Learning on Text Data Without User Identifiers

2025-01-10 · Yucheng Ding, Yangwenjian Tan, Xiangyu Liu, Chaoyue Niu 외

In many practical natural language applications, user data are highly sensitive, requiring anonymous uploads of text data from mobile devices to the cloud without user identifiers. However, the absence of user identifier…

Language ModelingLanguage Modelling

The Language Demographics of Amazon Mechanical Turk

2014-01-01 · TACL 2014 1 · Ellie Pavlick, Matt Post, Ann Irvine, Dmitry Kachaev 외

We present a large scale study of the languages spoken by bilingual workers on Mechanical Turk (MTurk). We establish a methodology for determining the language skills of anonymous crowd workers that is more robust than s…

Machine TranslationTranslation

YNTP-100: A Benchmark for Your Next Token Prediction with 100 People

2025-10-16 · Shiyao Ding, Takayuki Ito arxiv

Large language models (LLMs) trained for general \textit{next-token prediction} often fail to generate responses that reflect how specific individuals communicate. Progress on personalized alignment is further limited by…

Response Generation

Fine-Grained User Profiling for Personalized Task Matching in Mobile Crowdsensing

2018-11-14 · Yang Shuo, Zheng Zhenzhe, Tang Shaojie, Wu Fan 외

In mobile crowdsensing, finding the best match between tasks and users is crucial to ensure both the quality and effectiveness of a crowdsensing system. Existing works usually assume a centralized task assignment by the …

Recommendation Systems

Personalized Classifier Ensemble Pruning Framework for Mobile Crowdsourcing

2017-01-25 · Shaowei Wang, Liusheng Huang, Pengzhan Wang, Hongli Xu 외

Ensemble learning has been widely employed by mobile applications, ranging from environmental sensing to activity recognitions. One of the fundamental issue in ensemble learning is the trade-off between classification ac…

Ensemble LearningEnsemble Pruning