paper-with-me

홈 › Papers

Trustworthy Actionable Perturbations

2024-05-18 · Jesse Friedbaum, Sudarshan Adiga, Ravi Tandon

Counterfactuals, or modified inputs that lead to a different outcome, are an important tool for understanding the logic used by machine learning classifiers and how to change an undesirable classification. Even if a counterfactual changes a classifier's decision, however, it may not affect the true underlying class probabilities, i.e. the counterfactual may act like an adversarial attack and ``fool'' the classifier. We propose a new framework for creating modified inputs that change the true underlying probabilities in a beneficial way which we call Trustworthy Actionable Perturbations (TAP). This includes a novel verification procedure to ensure that TAP change the true class probabilities instead of acting adversarially. Our framework also includes new cost, reward, and goal definitions that are better suited to effectuating change in the real world. We present PAC-learnability results for our verification procedure and theoretically analyze our new method for measuring reward. We also develop a methodology for creating TAP and compare our results to those achieved by previous counterfactual methods.

📄 PDF Abstract BibTeX arXiv:2405.11195

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial Attackcounterfactual

Similar Papers 제목 키워드 기반

Engineering Trustworthy AI: A Developer Guide for Empirical Risk Minimization

2024-10-25 · Diana Pfau, Alexander Jung

AI systems increasingly shape critical decisions across personal and societal domains. While empirical risk minimization (ERM) drives much of the AI success, it typically prioritizes accuracy over trustworthiness, often …

Intelligent Knowledge Mining Framework: Bridging AI Analysis and Trustworthy Preservation

2025-12-19 · Binh Vu arxiv

The unprecedented proliferation of digital data presents significant challenges in access, integration, and value creation across all data-intensive sectors. Valuable information is frequently encapsulated within dispara…

RE-centric Recommendations for the Development of Trustworthy(er) Autonomous Systems

2023-05-29 · Krishna Ronanki, Beatriz Cabrero-Daniel, Jennifer Horkoff, Christian Berger

Complying with the EU AI Act (AIA) guidelines while developing and implementing AI systems will soon be mandatory within the EU. However, practitioners lack actionable instructions to operationalise ethics during AI syst…

Ethics

A Dempster-Shafer approach to trustworthy AI with application to fetal brain MRI segmentation

2022-04-05 · Lucas Fidon, Michael Aertsen, Florian Kofler, Andrea Bink 외

Deep learning models for medical image segmentation can fail unexpectedly and spectacularly for pathological cases and images acquired at different centers than training images, with labeling errors that violate expert k…

Image SegmentationMedical Image SegmentationMRI segmentationSemantic Segmentation

Vision Language Models Map Logos to Text via Semantic Entanglement in the Visual Projector

2025-10-14 · Sifan Li, Hongkai Chen, Yujun Cai, Qingwen Ye 외 arxiv

Vision Language Models (VLMs) have achieved impressive progress in multimodal reasoning; yet, they remain vulnerable to hallucinations, where outputs are not grounded in visual evidence. In this paper, we investigate a p…

Multimodal Reasoning