Shielded Representations: Protecting Sensitive Attributes Through Iterative Gradient-Based Projection
Natural language processing models tend to learn and encode social biases present in the data. One popular approach for addressing such biases is to eliminate encoded information from the model's representations. However, current methods are restricted to removing only linearly encoded information. In this work, we propose Iterative Gradient-Based Projection (IGBP), a novel method for removing non-linear encoded concepts from neural representations. Our method consists of iteratively training neural classifiers to predict a particular attribute we seek to eliminate, followed by a projection of the representation on a hypersurface, such that the classifiers become oblivious to the target attribute. We evaluate the effectiveness of our method on the task of removing gender and race information as sensitive attributes. Our results demonstrate that IGBP is effective in mitigating bias through intrinsic and extrinsic evaluations, with minimal impact on downstream task accuracy.
Code (1)
Tasks
AttributeSimilar Papers 제목 키워드 기반
Protecting gender and identity with disentangled speech representations
Besides its linguistic content, our speech is rich in biometric information that can be inferred by classifiers. Learning privacy-preserving representations for speech signals enables downstream tasks without sharing unn…
Privacy PreservingRepresentation LearningSpeaker VerificationSpeech RecognitionDeclarative Privacy-Preserving Inference Queries
Detecting inference queries running over personal attributes and protecting such queries from leaking individual information requires tremendous effort from practitioners. To tackle this problem, we propose an end-to-end…
Federated LearningManagementPrivacy PreservingPaCo-VLA: Passivity-Shielded Compliance Prior for Contact-Rich Vision-Language-Action Manipulation
Contact-rich manipulation demands both high-level semantic reasoning and the safe regulation of high-frequency contact dynamics. While Vision-Language-Action (VLA) models provide unprecedented semantic generalization, th…
Information Obfuscation of Graph Neural Networks
While the advent of Graph Neural Networks (GNNs) has greatly improved node and graph representation learning in many applications, the neighborhood aggregation scheme exposes additional vulnerabilities to adversaries see…
Adversarial DefenseGraph Representation LearningKnowledge GraphsRecommendation Systems+1Debiasing Diffusion Model: Enhancing Fairness through Latent Representation Learning in Stable Diffusion Model
Image generative models, particularly diffusion-based models, have surged in popularity due to their remarkable ability to synthesize highly realistic images. However, since these models are data-driven, they inherit bia…
FairnessmodelRepresentation Learning