paper-with-me

Papers

Will the Prince Get True Love's Kiss? On the Model Sensitivity to Gender Perturbation over Fairytale Texts

2023-10-16 · Christina Chance, Da Yin, Dakuo Wang, Kai-Wei Chang

Recent studies show that traditional fairytales are rife with harmful gender biases. To help mitigate these gender biases in fairytales, this work aims to assess learned biases of language models by evaluating their robustness against gender perturbations. Specifically, we focus on Question Answering (QA) tasks in fairytales. Using counterfactual data augmentation to the FairytaleQA dataset, we evaluate model robustness against swapped gender character information, and then mitigate learned biases by introducing counterfactual gender stereotypes during training time. We additionally introduce a novel approach that utilizes the massive vocabulary of language models to support text genres beyond fairytales. Our experimental results suggest that models are sensitive to gender perturbations, with significant performance drops compared to the original testing set. However, when first fine-tuned on a counterfactual training dataset, models are less sensitive to the later introduced anti-gender stereotyped text.

📄 PDF Abstract BibTeX arXiv:2310.10865

Code (0)

등록된 구현이 없습니다.

Tasks

counterfactualData AugmentationQuestion AnsweringSensitivity

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

KISS-Matcher: Fast and Robust Point Cloud Registration Revisited

2024-09-23 · Hyungtae Lim, Daebeom Kim, Gunhee Shin, Jingnan Shi 외

While global point cloud registration systems have advanced significantly in all aspects, many studies have focused on specific components, such as feature extraction, graph-theoretic pruning, or pose solvers. In this pa…

Point Cloud Registration

CLOVER: Closed-Loop Value Estimation and Ranking for End-to-End Autonomous Driving Planning

2026-05-14 · Sining Ang, Yuguang Yang, Canyu Chen, Yan Wang arxiv

End-to-end autonomous driving planners are commonly trained by imitating a single logged trajectory, yet evaluated by rule-based planning metrics that measure safety, feasibility, progress, and comfort. This creates a tr…

Autonomous Driving

Predicting Word Association Strengths

2017-09-01 · EMNLP 2017 9 · Andrew Cattle, Xiaojuan Ma

This paper looks at the task of predicting word association strengths across three datasets; WordNet Evocation (Boyd-Graber et al., 2006), University of Southern Florida Free Association norms (Nelson et al., 2004), and …

Word Embeddings

PRINCE: Provider-side Interpretability with Counterfactual Explanations in Recommender Systems

2019-11-19 · Azin Ghazimatin, Oana Balalau, Rishiraj Saha Roy, Gerhard Weikum

Interpretable explanations for recommender systems and other machine learning models are crucial to gain user trust. Prior works that have focused on paths connecting users and items in a heterogeneous network have sever…

counterfactualRecommendation Systems

Detecting Kissing Scenes in a Database of Hollywood Films

2019-06-05 · Amir Ziai

Detecting scene types in a movie can be very useful for application such as video editing, ratings assignment, and personalization. We propose a system for detecting kissing scenes in a movie. This system consists of two…

Kiss DetectionVideo Editing