paper-with-me

홈 › Papers

Confidential Guardian: Cryptographically Prohibiting the Abuse of Model Abstention

2025-05-29 · Stephan Rabanser, Ali Shahin Shamsabadi, Olive Franzese, Xiao Wang, Adrian Weller, Nicolas Papernot

Cautious predictions -- where a machine learning model abstains when uncertain -- are crucial for limiting harmful errors in safety-critical applications. In this work, we identify a novel threat: a dishonest institution can exploit these mechanisms to discriminate or unjustly deny services under the guise of uncertainty. We demonstrate the practicality of this threat by introducing an uncertainty-inducing attack called Mirage, which deliberately reduces confidence in targeted input regions, thereby covertly disadvantaging specific individuals. At the same time, Mirage maintains high predictive performance across all data points. To counter this threat, we propose Confidential Guardian, a framework that analyzes calibration metrics on a reference dataset to detect artificially suppressed confidence. Additionally, it employs zero-knowledge proofs of verified inference to ensure that reported confidence scores genuinely originate from the deployed model. This prevents the provider from fabricating arbitrary model confidence values while protecting the model's proprietary details. Our results confirm that Confidential Guardian effectively prevents the misuse of cautious predictions, providing verifiable assurances that abstention reflects genuine model uncertainty rather than malicious intent.

📄 PDF Abstract BibTeX arXiv:2505.23968

Code (1)

cleverhans-lab/confidential-guardian 공식 구현

Similar Papers 제목 키워드 기반

LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice

2025-01-19 · M. Mikail Demir, Hakan T. Otal, M. Abdullah Canbaz

Large Language Models (LLMs) hold promise for advancing legal practice by automating complex tasks and improving access to justice. However, their adoption is limited by concerns over client confidentiality, especially w…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+3

Privacy-Preserving Heterogeneous Federated Learning for Sensitive Healthcare Data

2024-06-15 · Yukai Xu, Jingfeng Zhang, Yujie Gu

In the realm of healthcare where decentralized facilities are prevalent, machine learning faces two major challenges concerning the protection of data and models. The data-level challenge concerns the data privacy leakag…

Federated LearningPrivacy Preserving

Model-Guardian: Protecting against Data-Free Model Stealing Using Gradient Representations and Deceptive Predictions

2025-03-23 · Yunfei Yang, Xiaojun Chen, Yuexin Xuan, Zhendong Zhao

Model stealing attack is increasingly threatening the confidentiality of machine learning models deployed in the cloud. Recent studies reveal that adversaries can exploit data synthesis techniques to steal machine learni…

model

Guarding the Guardians: Automated Analysis of Online Child Sexual Abuse

2023-08-07 · Juanita Puentes, Angela Castillo, Wilmar Osejo, Yuly Calderón 외

Online violence against children has increased globally recently, demanding urgent attention. Competent authorities manually analyze abuse complaints to comprehend crime dynamics and identify patterns. However, the manua…

Toward cryptographically verifiable authorization for autonomous AI agents: A security hypothesis, preliminary formal model, and proof-of-concept implementation

2026-07-23 · M. Llambí-Morillas, D. Fernández-Fernández arxiv

Autonomous AI agents increasingly execute actions, invoke tools, and operate on protected resources with limited human oversight. Existing authentication and authorization mechanisms establish identity and delegate autho…