paper-with-me

홈 › Papers

Adversarial Scrubbing of Demographic Information for Text Classification

2021-09-17 · EMNLP 2021 11 · Somnath Basu Roy Chowdhury, Sayan Ghosh, Yiyuan Li, Junier B. Oliva, Shashank Srivastava, Snigdha Chaturvedi

Contextual representations learned by language models can often encode undesirable attributes, like demographic associations of the users, while being trained for an unrelated target task. We aim to scrub such undesirable attributes and learn fair representations while maintaining performance on the target task. In this paper, we present an adversarial learning framework "Adversarial Scrubber" (ADS), to debias contextual representations. We perform theoretical analysis to show that our framework converges without leaking demographic information under certain conditions. We extend previous evaluation techniques by evaluating debiasing performance using Minimum Description Length (MDL) probing. Experimental evaluations on 8 datasets show that ADS generates representations with minimal information about demographic attributes while being maximally informative about the target task.

📄 PDF Abstract BibTeX arXiv:2109.08613

Code (1)

brcsomnath/adversarial-scrubber 공식 구현 pytorch

Tasks

Classificationtext-classificationText Classification

Similar Papers 제목 키워드 기반

Adversarial Removal of Demographic Attributes from Text Data

2018-08-20 · EMNLP 2018 10 · Yanai Elazar, Yoav Goldberg

Recent advances in Representation Learning and Adversarial Training seem to succeed in removing unwanted features from the learned representation. We show that demographic information of authors is encoded in -- and can …

Representation Learning

Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks

2025-07-08 · Huanming Shen, Baizhou Huang, Xiaojun Wan arxiv

Watermarking is a promising defense against the misuse of large language models (LLMs), yet it remains vulnerable to scrubbing and spoofing attacks. This vulnerability stems from an inherent trade-off governed by waterma…

$B^4$: A Black-Box Scrubbing Attack on LLM Watermarks

2024-11-02 · Baizhou Huang, Xiao Pu, Xiaojun Wan

Watermarking has emerged as a prominent technique for LLM-generated content detection by embedding imperceptible patterns. Despite supreme performance, its robustness against adversarial attacks remains underexplored. Pr…

Neural User Factor Adaptation for Text Classification: Learning to Generalize Across Author Demographics

2019-06-01 · SEMEVAL 2019 6 · Xiaolei Huang, Michael J. Paul

Language use varies across different demographic factors, such as gender, age, and geographic location. However, most existing document classification methods ignore demographic variability. In this study, we examine emp…

ClassificationDocument ClassificationGeneral Classificationtext-classification+1

Enterprise Disk Drive Scrubbing Based on Mondrian Conformal Predictors

2023-06-01 · Rahul Vishwakarma, Jinha Hwang, Soundouss Messoudi, Ava Hedayatipour

Disk scrubbing is a process aimed at resolving read errors on disks by reading data from the disk. However, scrubbing the entire storage array at once can adversely impact system performance, particularly during periods …

Conformal Prediction