paper-with-me

홈 › Papers

Grader variability and the importance of reference standards for evaluating machine learning models for diabetic retinopathy

2017-10-04 · Jonathan Krause, Varun Gulshan, Ehsan Rahimy, Peter Karth, Kasumi Widner, Greg S. Corrado, Lily Peng, Dale R. Webster

Diabetic retinopathy (DR) and diabetic macular edema are common complications of diabetes which can lead to vision loss. The grading of DR is a fairly complex process that requires the detection of fine features such as microaneurysms, intraretinal hemorrhages, and intraretinal microvascular abnormalities. Because of this, there can be a fair amount of grader variability. There are different methods of obtaining the reference standard and resolving disagreements between graders, and while it is usually accepted that adjudication until full consensus will yield the best reference standard, the difference between various methods of resolving disagreements has not been examined extensively. In this study, we examine the variability in different methods of grading, definitions of reference standards, and their effects on building deep learning models for the detection of diabetic eye disease. We find that a small set of adjudicated DR grades allows substantial improvements in algorithm performance. The resulting algorithm's performance was on par with that of individual U.S. board-certified ophthalmologists and retinal specialists.

📄 PDF Abstract BibTeX arXiv:1710.01711

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine Learning

Similar Papers 제목 키워드 기반

Automated segmentation of choroidal layers from 3-dimensional macular optical coherence tomography scans

2021-03-11 · Kyungmoo Lee, Alexis K. Warren, Michael D. Abramoff, Andreas Wahle 외

Background: Changes in choroidal thickness are associated with various ocular diseases and the choroid can be imaged using spectral-domain optical coherence tomography (SDOCT) and enhanced depth imaging OCT (EDIOCT). New…

Segmentation

Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering

2026-05-08 · Zewei Tian, Alex Liu, Lief Esbenshade, Michael Xiao 외 arxiv

The integration of large language models (LLMs) into educational assessment represents a transformative shift in classroom grading practices. While automated scoring systems and machine learning techniques have existed f…

Prompt Engineering

Graders should cheat: privileged information enables expert-level automated evaluations

2025-02-16 · Jin Peng Zhou, Sébastien M. R. Arnold, Nan Ding, Kilian Q. Weinberger 외

Auto-evaluating language models (LMs), i.e., using a grader LM to evaluate the candidate LM, is an appealing way to accelerate the evaluation process and the cost associated with it. But this presents a paradox: how can …

Math

EvoGrader: an online formative assessment tool for automatically evaluating written evolutionary explanations

2016-01-13 · Kayhan Moharreri, Minsu Ha, Ross H Nehm

EvoGrader is a free, online, on-demand formative assessment service designed for use in undergraduate biology classrooms. EvoGrader's web portal is powered by Amazon's Elastic Cloud and run with LightSIDE Lab's open-sour…

BIG-bench Machine Learning

Skewed Score: A statistical framework to assess autograders

2025-07-04 · Magda Dubois, Harry Coppock, Mario Giulianelli, Timo Flesch 외 arxiv

The evaluation of large language model (LLM) outputs is increasingly performed by other LLMs, a setup commonly known as "LLM-as-a-judge", or autograders. While autograders offer a scalable alternative to human evaluation…

Bias Detection