Adversarial Scrubbing of Demographic Information for Text Classification
Contextual representations learned by language models can often encode undesirable attributes, like demographic associations of the users, while being trained for an unrelated target task. We aim to scrub such undesirable attributes and learn fair representations while maintaining performance on the target task. In this paper, we present an adversarial learning framework "Adversarial Scrubber" (ADS), to debias contextual representations. We perform theoretical analysis to show that our framework converges without leaking demographic information under certain conditions. We extend previous evaluation techniques by evaluating debiasing performance using Minimum Description Length (MDL) probing. Experimental evaluations on 8 datasets show that ADS generates representations with minimal information about demographic attributes while being maximally informative about the target task.
Code (1)
Tasks
Classificationtext-classificationText ClassificationSimilar Papers 제목 키워드 기반
Adversarial Removal of Demographic Attributes from Text Data
Recent advances in Representation Learning and Adversarial Training seem to succeed in removing unwanted features from the learned representation. We show that demographic information of authors is encoded in -- and can …
Representation LearningEnhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks
Watermarking is a promising defense against the misuse of large language models (LLMs), yet it remains vulnerable to scrubbing and spoofing attacks. This vulnerability stems from an inherent trade-off governed by waterma…
$B^4$: A Black-Box Scrubbing Attack on LLM Watermarks
Watermarking has emerged as a prominent technique for LLM-generated content detection by embedding imperceptible patterns. Despite supreme performance, its robustness against adversarial attacks remains underexplored. Pr…
Neural User Factor Adaptation for Text Classification: Learning to Generalize Across Author Demographics
Language use varies across different demographic factors, such as gender, age, and geographic location. However, most existing document classification methods ignore demographic variability. In this study, we examine emp…
ClassificationDocument ClassificationGeneral Classificationtext-classification+1Enterprise Disk Drive Scrubbing Based on Mondrian Conformal Predictors
Disk scrubbing is a process aimed at resolving read errors on disks by reading data from the disk. However, scrubbing the entire storage array at once can adversely impact system performance, particularly during periods …
Conformal Prediction