paper-with-me

홈 › Papers

Responsible AI for Test Equity and Quality: The Duolingo English Test as a Case Study

2024-08-28 · Jill Burstein, Geoffrey T. LaFlair, Kevin Yancey, Alina A. von Davier, Ravit Dotan

Artificial intelligence (AI) creates opportunities for assessments, such as efficiencies for item generation and scoring of spoken and written responses. At the same time, it poses risks (such as bias in AI-generated item content). Responsible AI (RAI) practices aim to mitigate risks associated with AI. This chapter addresses the critical role of RAI practices in achieving test quality (appropriateness of test score inferences), and test equity (fairness to all test takers). To illustrate, the chapter presents a case study using the Duolingo English Test (DET), an AI-powered, high-stakes English language assessment. The chapter discusses the DET RAI standards, their development and their relationship to domain-agnostic RAI principles. Further, it provides examples of specific RAI practices, showing how these practices meaningfully address the ethical principles of validity and reliability, fairness, privacy and security, and transparency and accountability standards to ensure test equity and quality.

📄 PDF Abstract BibTeX arXiv:2409.07476

Code (0)

등록된 구현이 없습니다.

Tasks

Fairness

Similar Papers 제목 키워드 기반

Context Based Approach for Second Language Acquisition

2018-06-01 · WS 2018 6 · Nihal V. Nayak, Arjun R. Rao

SLAM 2018 focuses on predicting a student{'}s mistake while using the Duolingo application. In this paper, we describe the system we developed for this shared task. Our system uses a logistic regression model to predict …

es-enfr-enLanguage Acquisition

Training and Inference Methods for High-Coverage Neural Machine Translation

2020-07-01 · WS 2020 7 · Michael Yang, Yixin Liu, Rahul Mayuranath

In this paper, we introduce a system built for the Duolingo Simultaneous Translation And Paraphrase for Language Education (STAPLE) shared task at the 4th Workshop on Neural Generation and Translation (WNGT 2020). We par…

DiversityMachine TranslationTranslationVocal Bursts Intensity Prediction

Jump-Starting Item Parameters for Adaptive Language Tests

2021-11-01 · EMNLP 2021 11 · Arya D. McCarthy, Kevin P. Yancey, Geoff T. LaFlair, Jesse Egbert 외

A challenge in designing high-stakes language assessments is calibrating the test item difficulties, either a priori or from limited pilot test data. While prior work has addressed ‘cold start’ estimation of item difficu…

Language AcquisitionMulti-Task LearningSkills Assessment

Machine Learning--Driven Language Assessment

2020-01-01 · TACL 2020 1 · Burr Settles, Geoffrey T. LaFlair, Masato Hagiwara

We describe a method for rapidly creating language proficiency assessments, and provide experimental evidence that such tests can be valid, reliable, and secure. Our approach is the first to use machine learning and natu…

BIG-bench Machine LearningLanguage AcquisitionSkills Assessmentvalid

POSTECH Submission on Duolingo Shared Task

2020-07-01 · WS 2020 7 · Junsu Park, Hong-Seok Kwon, Jong-Hyeok Lee

In this paper, we propose a transfer learning based simultaneous translation model by extending BART. We pre-trained BART with Korean Wikipedia and a Korean news dataset, and fine-tuned with an additional web-crawled par…

Transfer LearningTranslation