paper-with-me

홈 › Papers

Potential and Perils of Large Language Models as Judges of Unstructured Textual Data

2025-01-14 · Rewina Bedemariam, Natalie Perez, Sreyoshi Bhaduri, Satya Kapoor, Alex Gil, Elizabeth Conjar, Ikkei Itoku, David Theil, Aman Chadha, Naumaan Nayyar

Rapid advancements in large language models have unlocked remarkable capabilities when it comes to processing and summarizing unstructured text data. This has implications for the analysis of rich, open-ended datasets, such as survey responses, where LLMs hold the promise of efficiently distilling key themes and sentiments. However, as organizations increasingly turn to these powerful AI systems to make sense of textual feedback, a critical question arises, can we trust LLMs to accurately represent the perspectives contained within these text based datasets? While LLMs excel at generating human-like summaries, there is a risk that their outputs may inadvertently diverge from the true substance of the original responses. Discrepancies between the LLM-generated outputs and the actual themes present in the data could lead to flawed decision-making, with far-reaching consequences for organizations. This research investigates the effectiveness of LLM-as-judge models to evaluate the thematic alignment of summaries generated by other LLMs. We utilized an Anthropic Claude model to generate thematic summaries from open-ended survey responses, with Amazon's Titan Express, Nova Pro, and Meta's Llama serving as judges. This LLM-as-judge approach was compared to human evaluations using Cohen's kappa, Spearman's rho, and Krippendorff's alpha, validating a scalable alternative to traditional human centric evaluation methods. Our findings reveal that while LLM-as-judge offer a scalable solution comparable to human raters, humans may still excel at detecting subtle, context-specific nuances. Our research contributes to the growing body of knowledge on AI assisted text analysis. Further, we provide recommendations for future research, emphasizing the need for careful consideration when generalizing LLM-as-judge models across various contexts and use cases.

📄 PDF Abstract BibTeX arXiv:2501.08167

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…

Similar Papers 제목 키워드 기반

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

2024-12-07 · Haitao Li, Qian Dong, Junjie Chen, Huixue Su 외

The rapid advancement of Large Language Models (LLMs) has driven their expanding application across various fields. One of the most promising applications is their role as evaluators based on natural language responses, …

Curse of Knowledge: When Complex Evaluation Context Benefits yet Biases LLM Judges

2025-09-03 · Weiyuan Li, Xintao Wang, Siyu Yuan, Rui Xu 외 arxiv

As large language models (LLMs) grow more capable, they face increasingly diverse and complex tasks, making reliable evaluation challenging. The paradigm of LLMs as judges has emerged as a scalable solution, yet prior wo…

From Perils to Possibilities: Understanding how Human (and AI) Biases affect Online Fora

2024-03-21 · Virginia Morini, Valentina Pansanella, Katherine Abramski, Erica Cau 외

Social media platforms are online fora where users engage in discussions, share content, and build connections. This review explores the dynamics of social interactions, user-generated contents, and biases within the con…

Misinformation

Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation

2025-05-22 · Jiwon Moon, Yerin Hwang, Dongryeol Lee, Taegwan Kang 외

With the growing use of large language models(LLMs) as evaluators, their application has expanded to code evaluation tasks, where they assess the correctness of generated code without relying on reference implementations…

Humans or LLMs as the Judge? A Study on Judgement Biases

2024-02-16 · Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang 외

Adopting human and large language models (LLM) as judges (a.k.a human- and LLM-as-a-judge) for evaluating the performance of LLMs has recently gained attention. Nonetheless, this approach concurrently introduces potentia…

Misinformation