paper-with-me

홈 › Papers

Information Theoretic Evaluation of Privacy-Leakage, Interpretability, and Transferability for Trustworthy AI

2021-06-06 · Mohit Kumar, Bernhard A. Moser, Lukas Fischer, Bernhard Freudenthaler

In order to develop machine learning and deep learning models that take into account the guidelines and principles of trustworthy AI, a novel information theoretic trustworthy AI framework is introduced. A unified approach to "privacy-preserving interpretable and transferable learning" is considered for studying and optimizing the tradeoffs between privacy, interpretability, and transferability aspects. A variational membership-mapping Bayesian model is used for the analytical approximations of the defined information theoretic measures for privacy-leakage, interpretability, and transferability. The approach consists of approximating the information theoretic measures via maximizing a lower-bound using variational optimization. The study presents a unified information theoretic approach to study different aspects of trustworthy AI in a rigorous analytical manner. The approach is demonstrated through numerous experiments on benchmark datasets and a real-world biomedical application concerned with the detection of mental stress on individuals using heart rate variability analysis.

📄 PDF Abstract BibTeX arXiv:2106.06046

Code (0)

등록된 구현이 없습니다.

Tasks

Heart Rate VariabilityPrivacy Preserving

Similar Papers 제목 키워드 기반

Beyond Verification: Abductive Explanations for Post-AI Assessment of Privacy Leakage

2025-11-13 · Belona Sonna, Alban Grastien, Claire Benn arxiv

Privacy leakage in AI-based decision processes poses significant risks, particularly when sensitive information can be inferred. We propose a formal framework to audit privacy leakage using abductive explanations, which …

Membership Inference Attacks fueled by Few-Short Learning to detect privacy leakage tackling data integrity

2025-03-12 · Daniel Jiménez-López, Nuria Rodríguez-Barroso, M. Victoria Luzón, Francisco Herrera

Deep learning models have an intrinsic privacy issue as they memorize parts of their training data, creating a privacy leakage. Membership Inference Attacks (MIA) exploit it to obtain confidential information about the d…

Deep LearningFew-Shot Learningimage-classificationImage Classification+2

PrivacyScalpel: Enhancing LLM Privacy via Interpretable Feature Intervention with Sparse Autoencoders

2025-03-14 · Ahmed Frikha, Muhammad Reza Ar Razi, Krishna Kanth Nakka, Ricardo Mendes 외

Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language processing but also pose significant privacy risks by memorizing and leaking Personally Identifiable Information (PII). Existing …

MemorizationPrivacy Preserving

Balancing Privacy Protection and Interpretability in Federated Learning

2023-02-16 · Zhe Li, Honglong Chen, Zhichen Ni, Huajie Shao

Federated learning (FL) aims to collaboratively train the global model in a distributed manner by sharing the model parameters from local clients to a central server, thereby potentially protecting users' private informa…

Federated Learning

Chain-of-Sanitized-Thoughts: Plugging PII Leakage in CoT of Large Reasoning Models

2026-01-08 · Arghyadeep Das, Sai Sreenivas Chintha, Rishiraj Girmal, Kinjal Pandey 외 arxiv

Large Reasoning Models (LRMs) improve performance, reliability, and interpretability by generating explicit chain-of-thought (CoT) reasoning, but this transparency introduces a serious privacy risk: intermediate reasonin…