paper-with-me

Papers

Wizard of Errors: Introducing and Evaluating Machine Learning Errors in Wizard of Oz Studies

2023-02-17 · Anniek Jansen, Sara Colombo

When designing Machine Learning (ML) enabled solutions, designers often need to simulate ML behavior through the Wizard of Oz (WoZ) approach to test the user experience before the ML model is available. Although reproducing ML errors is essential for having a good representation, they are rarely considered. We introduce Wizard of Errors (WoE), a tool for conducting WoZ studies on ML-enabled solutions that allows simulating ML errors during user experience assessment. We explored how this system can be used to simulate the behavior of a computer vision model. We tested WoE with design students to determine the importance of considering ML errors in design, the relevance of using descriptive error types instead of confusion matrix, and the suitability of manual error control in WoZ studies. Our work identifies several challenges, which prevent realistic error representation by designers in such studies. We discuss the implications of these findings for design.

📄 PDF Abstract BibTeX arXiv:2302.08799

Code (0)

등록된 구현이 없습니다.

Tasks

Descriptive

Methods 이 논문이 사용한 방법론

Wizard Computer vision is an interesting tool for animal behavior monitoring, mainly because it limits animal handling and it can be used to record various traits using only one sensor.…
Test 설명 없음

Similar Papers 제목 키워드 기반

Improvement on Extrapolation of Species Abundance Distribution Across Scales from Moments Across Scales

2020-07-01 · Saeid Alirezazadeh, Khadijeh Alibabaei

Raw moments are used as a way to estimate species abundance distribution. The almost linear pattern of the log transformation of raw moments across scales allow us to extrapolate species abundance distribution for larger…

Cipher: A Prototype Game-with-a-Purpose for Detecting Errors in Text

2020-05-01 · LREC 2020 5 · Liang Xu, Jon Chamberlain

Errors commonly exist in machine-generated documents and publication materials; however, some correction algorithms do not perform well for complex errors and it is costly to employ humans to do the task. To solve the pr…

The PHOTON Wizard -- Towards Educational Machine Learning Code Generators

2020-02-13 · Ramona Leenings, Nils Ralf Winter, Kelvin Sarink, Jan Ernsting 외

Despite the tremendous efforts to democratize machine learning, especially in applied-science, the application is still often hampered by the lack of coding skills. As we consider programmatic understanding key to buildi…

BIG-bench Machine Learningvalid

DaLA: Danish Linguistic Acceptability Evaluation Guided by Real World Errors

2025-12-04 · Gianluca Barmina, Nathalie Carmen Hau Norman, Peter Schneider-Kamp, Lukas Galke Poech arxiv

We present an enhanced benchmark for evaluating linguistic acceptability in Danish. We first analyze the most common errors found in written Danish. Based on this analysis, we introduce a set of fourteen corruption funct…

Linguistic Acceptability

MISMATCH: Fine-grained Evaluation of Machine-generated Text with Mismatch Error Types

2023-06-18 · Keerthiram Murugesan, Sarathkrishna Swaminathan, Soham Dan, Subhajit Chaudhury 외

With the growing interest in large language models, the need for evaluating the quality of machine text compared to reference (typically human-generated) text has become focal attention. Most recent works focus either on…

Sentence