paper-with-me

홈 › Papers

Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings

2023-08-03 · Veronika Hackl, Alexandra Elena Müller, Michael Granitzer, Maximilian Sailer

This study investigates the consistency of feedback ratings generated by OpenAI's GPT-4, a state-of-the-art artificial intelligence language model, across multiple iterations, time spans and stylistic variations. The model rated responses to tasks within the Higher Education (HE) subject domain of macroeconomics in terms of their content and style. Statistical analysis was conducted in order to learn more about the interrater reliability, consistency of the ratings across iterations and the correlation between ratings in terms of content and style. The results revealed a high interrater reliability with ICC scores ranging between 0.94 and 0.99 for different timespans, suggesting that GPT-4 is capable of generating consistent ratings across repetitions with a clear prompt. Style and content ratings show a high correlation of 0.87. When applying a non-adequate style the average content ratings remained constant, while style ratings decreased, which indicates that the large language model (LLM) effectively distinguishes between these two criteria during evaluation. The prompt used in this study is furthermore presented and explained. Further research is necessary to assess the robustness and reliability of AI models in various use cases.

📄 PDF Abstract BibTeX arXiv:2308.02575

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Model

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Adam 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Residual Connection 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

ACORN: Aspect-wise Commonsense Reasoning Explanation Evaluation

2024-05-08 · Ana Brassard, Benjamin Heinzerling, Keito Kudo, Keisuke Sakaguchi 외

Evaluating the quality of free-text explanations is a multifaceted, subjective, and labor-intensive task. Large language models (LLMs) present an appealing alternative due to their potential for consistency, scalability,…

From Text to Insight: Leveraging Large Language Models for Performance Evaluation in Management

2024-08-09 · Ning li, Huaikang Zhou, Mingze Xu

This study explores the potential of Large Language Models (LLMs), specifically GPT-4, to enhance objectivity in organizational task performance evaluations. Through comparative analyses across two studies, including var…

Management

Towards Best Experiment Design for Evaluating Dialogue System Output

2019-09-23 · WS 2019 10 · Sashank Santhanam, Samira Shaikh

To overcome the limitations of automated metrics (e.g. BLEU, METEOR) for evaluating dialogue systems, researchers typically use human judgments to provide convergent evidence. While it has been demonstrated that human ju…

Dialogue Evaluation

k-Rater Reliability: The Correct Unit of Reliability for Aggregated Human Annotations

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Since the inception of crowdsourcing, aggregation has been a common strategy for dealing with unreliable data. Aggregate ratings are more reliable than individual ones. However, many NLP datasets that rely on aggregate r…

Word Similarity

k-Rater Reliability: The Correct Unit of Reliability for Aggregated Human Annotations

2022-03-24 · ACL 2022 5 · Ka Wong, Praveen Paritosh

Since the inception of crowdsourcing, aggregation has been a common strategy for dealing with unreliable data. Aggregate ratings are more reliable than individual ones. However, many natural language processing (NLP) app…