paper-with-me

홈 › Papers

A Step Towards Mixture of Grader: Statistical Analysis of Existing Automatic Evaluation Metrics

2024-10-13 · Yun Joon Soh, Jishen Zhao

The explosion of open-sourced models and Question-Answering (QA) datasets emphasizes the importance of automated QA evaluation. We studied the statistics of the existing evaluation metrics for a better understanding of their limitations. By measuring the correlation coefficients of each evaluation metric concerning human-like evaluation score, we observed the following: (1) existing metrics have a high correlation among them concerning the question type (e.g., single word, single phrase, etc.), (2) no single metric can adequately estimate the human-like evaluation. As a potential solution, we discuss how a Mixture Of Grader could potentially improve the auto QA evaluator quality.

📄 PDF Abstract BibTeX arXiv:2410.10030

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Skewed Score: A statistical framework to assess autograders

2025-07-04 · Magda Dubois, Harry Coppock, Mario Giulianelli, Timo Flesch 외 arxiv

The evaluation of large language model (LLM) outputs is increasingly performed by other LLMs, a setup commonly known as "LLM-as-a-judge", or autograders. While autograders offer a scalable alternative to human evaluation…

Bias Detection

Assessment of L2 Oral Proficiency using Speech Large Language Models

2025-05-27 · Rao Ma, Mengjie Qian, Siyuan Tang, Stefano Bannò 외

The growing population of L2 English speakers has increased the demand for developing automatic graders for spoken language assessment (SLA). Historically, statistical models, text encoders, and self-supervised speech mo…

Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System

2026-06-23 · Tian Zheng, Kai-Tai Hsu arxiv

Agentic data analysis systems produce rich outputs, including code, numerical results, and verbal diagnostics. This makes them more challenging to evaluate than single-turn LLM responses. It is therefore necessary to dis…

Evaluating Complex Task through Crowdsourcing: Multiple Views Approach

2017-03-30 · Lingyu Lyu, Mehmed Kantardzic

With the popularity of massive open online courses, grading through crowdsourcing has become a prevalent approach towards large scale classes. However, for getting grades for complex tasks, which require specific skills …

ClipGrader: Leveraging Vision-Language Models for Robust Label Quality Assessment in Object Detection

2025-03-03 · Hong Lu, Yali Bian, Rahul C. Shah

High-quality annotations are essential for object detection models, but ensuring label accuracy - especially for bounding boxes - remains both challenging and costly. This paper introduces ClipGrader, a novel approach th…

Objectobject-detectionObject DetectionPseudo Label+1