paper-with-me

Papers

Measuring AI Systems Beyond Accuracy

2022-04-07 · Violet Turri, Rachel Dzombak, Eric Heim, Nathan VanHoudnos, Jay Palat, Anusha Sinha

Current test and evaluation (T&E) methods for assessing machine learning (ML) system performance often rely on incomplete metrics. Testing is additionally often siloed from the other phases of the ML system lifecycle. Research investigating cross-domain approaches to ML T&E is needed to drive the state of the art forward and to build an Artificial Intelligence (AI) engineering discipline. This paper advocates for a robust, integrated approach to testing by outlining six key questions for guiding a holistic T&E strategy.

📄 PDF Abstract BibTeX arXiv:2204.04211

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistency

2020-06-30 · NeurIPS 2020 12 · Robert Geirhos, Kristof Meding, Felix A. Wichmann

A central problem in cognitive science and behavioural neuroscience as well as in machine learning and artificial intelligence research is to ascertain whether two or more decision makers (be they brains or algorithms) u…

Decision MakingObject Recognition

LABBench2: An Improved Benchmark for AI Systems Performing Biology Research

2026-02-04 · Jon M Laurent, Albert Bou, Michael Pieler, Conor Igoe 외 arxiv

Optimism for accelerating scientific discovery with AI continues to grow. Current applications of AI in scientific research range from training dedicated foundation models on scientific data to agentic autonomous hypothe…

Measuring Intelligence Beyond Human Scale

2026-07-08 · Jerry Han, Rafael Moschopoulos, Ella Colby, Vishrut Goyal 외 arxiv

How can we measure intelligence beyond human capability? Human-authored benchmarks saturate, and above human capability, examiners may not know which tasks are both hard and verifiable. We argue that this difficulty is i…

Measuring Lexical Cohesion: Beyond Word Repetition

2014-08-01 · COLING 2014 8 · Anna Kazantseva, Stan Szpakowicz

Beyond Accuracy: Measuring Logical Compliance of Predictive Models

2026-06-18 · Guillaume Olivier Delplanque, Pierre Genevès, Nabil Layaïda, Zephirin Faure arxiv

Machine learning models are predominantly evaluated through predictive performance metrics such as ranking quality, prediction error, or classification accuracy. While these metrics effectively quantify how closely predi…

Link Prediction