paper-with-me

홈 › Papers

TIDE: Textual Identity Detection for Evaluating and Augmenting Classification and Language Models

2023-09-07 · Emmanuel Klu, Sameer Sethi

Machine learning models can perpetuate unintended biases from unfair and imbalanced datasets. Evaluating and debiasing these datasets and models is especially hard in text datasets where sensitive attributes such as race, gender, and sexual orientation may not be available. When these models are deployed into society, they can lead to unfair outcomes for historically underrepresented groups. In this paper, we present a dataset coupled with an approach to improve text fairness in classifiers and language models. We create a new, more comprehensive identity lexicon, TIDAL, which includes 15,123 identity terms and associated sense context across three demographic categories. We leverage TIDAL to develop an identity annotation and augmentation tool that can be used to improve the availability of identity context and the effectiveness of ML fairness techniques. We evaluate our approaches using human contributors, and additionally run experiments focused on dataset and model debiasing. Results show our assistive annotation technique improves the reliability and velocity of human-in-the-loop processes. Our dataset and methods uncover more disparities during evaluation, and also produce more fair models during remediation. These approaches provide a practical path forward for scaling classifier and generative model fairness in real-world settings.

📄 PDF Abstract BibTeX arXiv:2309.04027

Code (0)

등록된 구현이 없습니다.

Tasks

Fairness

Similar Papers 제목 키워드 기반

TIDE: Achieving Balanced Subject-Driven Image Generation via Target-Instructed Diffusion Enhancement

2025-09-08 · Jibai Lin, Bo Ma, Yating Yang, Xi Zhou 외 arxiv

Subject-driven image generation (SDIG) aims to manipulate specific subjects within images while adhering to textual instructions, a task crucial for advancing text-to-image diffusion models. SDIG requires reconciling the…

Image Generation

TIDE: Every Layer Knows the Token Beneath the Context

2026-05-07 · Ajay Jaiswal, Lauren Hannah, Han-Byul Kim, Duc Hoang 외 arxiv

We revisit a universally accepted but under-examined design choice in every modern LLM: a token index is looked up once at the input embedding layer and then permanently discarded. This single-injection assumption induce…

Peptides of H. sapiens and P. falciparum that are predicted to bind strongly to HLA-A*24:02 and homologous to a SARS-CoV-2 peptide

2021-01-18 · Yekbun Adiguzel

Aim: This study is looking for a common pathogenicity between SARS-CoV-2 and plasmodium species, in individuals with certain HLA serotypes. Methods: 1-) Tblastx searches of SARS-CoV-2 are performed by limiting searches t…

Improving Protein-peptide Interface Predictions in the Low Data Regime

2023-05-31 · Justin Diamond, Markus Lill

We propose a novel approach for predicting protein-peptide interactions using a bi-modal transformer architecture that learns an inter-facial joint distribution of residual contacts. The current data sets for crystallize…

Towards Robust Argumentative Essay Understanding via TIDE: An Interactive Framework with Trial and Debate

2026-05-17 · Zheqin Yin, Yupei Ren, Yadong Zhang, Yujiang Lu 외 arxiv

Argumentative essays serve as a vital medium for assessing critical thinking and reasoning skills, yet there is limited works on accurately understanding and evaluating such texts via prompt. In this work, we propose TID…

Automated Essay Scoring