paper-with-me

Papers

Multisource AI Scorecard Table for System Evaluation

2021-02-08 · Erik Blasch, James Sung, Tao Nguyen

The paper describes a Multisource AI Scorecard Table (MAST) that provides the developer and user of an artificial intelligence (AI)/machine learning (ML) system with a standard checklist focused on the principles of good analysis adopted by the intelligence community (IC) to help promote the development of more understandable systems and engender trust in AI outputs. Such a scorecard enables a transparent, consistent, and meaningful understanding of AI tools applied for commercial and government use. A standard is built on compliance and agreement through policy, which requires buy-in from the stakeholders. While consistency for testing might only exist across a standard data set, the community requires discussion on verification and validation approaches which can lead to interpretability, explainability, and proper use. The paper explores how the analytic tradecraft standards outlined in Intelligence Community Directive (ICD) 203 can provide a framework for assessing the performance of an AI system supporting various operational needs. These include sourcing, uncertainty, consistency, accuracy, and visualization. Three use cases are presented as notional examples that support security for comparative analysis.

📄 PDF Abstract BibTeX arXiv:2102.03985

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fighting Sampling Bias: A Framework for Training and Evaluating Credit Scoring Models

2024-07-17 · Nikita Kozodoi, Stefan Lessmann, Morteza Alamgir, Luis Moreira-Matias 외

Scoring models support decision-making in financial institutions. Their estimation and evaluation are based on the data of previously accepted applicants with known repayment behavior. This creates sampling bias: the ava…

Self-Learning

AI Data Development: A Scorecard for the System Card Framework

2025-06-02 · Tadesse K. Bahiru, Haileleol Tibebu, Ioannis A. Kakadiaris

Artificial intelligence has transformed numerous industries, from healthcare to finance, enhancing decision-making through automated systems. However, the reliability of these systems is mainly dependent on the quality o…

Fairness

Mind the (Language) Gap: Towards Probing Numerical and Cross-Lingual Limits of LVLMs

2025-08-24 · Somraj Gautam, Abhirama Subramanyam Penamakuri, Abhishek Bhandari, Gaurav Harit arxiv

We introduce MMCRICBENCH-3K, a benchmark for Visual Question Answering (VQA) on cricket scorecards, designed to evaluate large vision-language models (LVLMs) on complex numerical and cross-lingual reasoning over semi-str…

Visual Question Answering

A Vertical Federated Learning Method for Interpretable Scorecard and Its Application in Credit Scoring

2020-09-14 · Fanglan Zheng, Erihe, Kun Li, Jiang Tian 외

With the success of big data and artificial intelligence in many fields, the applications of big data driven models are expected in financial risk management especially credit scoring and rating. Under the premise of dat…

Federated LearningManagementVertical Federated Learning

Fairness in Credit Scoring: Assessment, Implementation and Profit Implications

2021-03-02 · Nikita Kozodoi, Johannes Jacob, Stefan Lessmann

The rise of algorithmic decision-making has spawned much research on fair machine learning (ML). Financial institutions use ML for building risk scorecards that support a range of credit-related decisions. Yet, the liter…

Decision MakingFairness