The Viability of Crowdsourcing for RAG Evaluation
How good are humans at writing and judging responses in retrieval-augmented generation (RAG) scenarios? To answer this question, we investigate the efficacy of crowdsourcing for RAG through two complementary studies: response writing and response utility judgment. We present the Crowd RAG Corpus 2025 (CrowdRAG-25), which consists of 903 human-written and 903 LLM-generated responses for the 301 topics of the TREC RAG'24 track, across the three discourse styles 'bulleted list', 'essay', and 'news'. For a selection of 65 topics, the corpus further contains 47,320 pairwise human judgments and 10,556 pairwise LLM judgments across seven utility dimensions (e.g., coverage and coherence). Our analyses give insights into human writing behavior for RAG and the viability of crowdsourcing for RAG evaluation. Human pairwise judgments provide reliable and cost-effective results compared to LLM-based pairwise or human/LLM-based pointwise judgments, as well as automated comparisons with human-written reference responses. All our data and tools are freely available.
Code (1)
Tasks
RAGRetrieval-augmented GenerationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A Probabilistic Annotation Model for Crowdsourcing Coreference
The availability of large scale annotated corpora for coreference is essential to the development of the field. However, creating resources at the required scale via expert annotation would be too expensive. Crowdsourcin…
Coreference ResolutionmodelQuestion AnsweringDesigning Search Tasks for Archive Search
Longitudinal corpora like legal, corporate and newspaper archives are of immense value to a variety of users, and time as an important factor strongly influences their search behavior in these archives. While many system…
Moving outside the lab: The viability of conducting sensorimotor learning studies online
Collecting data online via crowdsourcing platforms has proven to be a very efficient way to recruit a large and diverse sample. Studies of motor learning, however, have been largely confined to the lab due to the need fo…
Crowdsourcing for Evaluating Machine Translation Quality
The recent popularity of machine translation has increased the demand for the evaluation of translations. However, the traditional evaluation approach, manual checking by a bilingual professional, is too expensive and to…
Machine TranslationSentenceTranslationOnline Domain Adaptation for Continuous Cross-Subject Liver Viability Evaluation Based on Irregular Thermal Data
Accurate evaluation of liver viability during its procurement is a challenging issue and has traditionally been addressed by taking invasive biopsy on liver. Recently, people have started to investigate on the non-invasi…
Domain AdaptationOnline Domain Adaptation