paper-with-me

Papers

Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes

2025-09-25 · Guangliang Liu, Bocheng Chen, Han Zi, Xitong Zhang, Kristen Marie Johnson arxiv

Moral alignment has emerged as a widely adopted approach for regulating the behavior of pretrained language models (PLMs), typically through fine-tuning on curated datasets. Gender stereotype mitigation is a representational task within the broader application of moral alignment. However, this process often comes at the cost of degraded downstream task performance. Prior studies commonly aim to achieve a performance trade-off by encouraging PLMs to selectively forget only stereotypical knowledge through carefully designed fairness objective, while preserving their language modeling capability (overall forgetting). In this short paper, we investigate whether the performance trade-off can be achieved through the lens of forgetting and the fairness objective. Our analysis shows that the large datasets needed for satisfactory fairness highlight the limitations of current fairness objectives in achieving an effective trade-off: (1) downstream task performance is strongly correlated with overall forgetting; (2) selective forgetting reduces stereotypes, but overall forgetting increases. and (3) general solutions for alleviating forgetting are ineffective at reducing the overall forgetting and fail to improve downstream task performance.

📄 PDF Abstract BibTeX arXiv:2509.21456

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Do VLMs Have a Moral Backbone? A Study on the Fragile Morality of Vision-Language Models

2026-01-23 · Zhining Liu, Tianyi Wang, Xiao Lin, Penghao Ouyang 외 arxiv

Despite substantial efforts toward improving the moral alignment of Vision-Language Models (VLMs), it remains unclear whether their ethical judgments are stable in realistic settings. This work studies moral robustness i…

The Straight and Narrow: Do LLMs Possess an Internal Moral Path?

2026-01-15 · Luoming Hu, Jingjie Zeng, Liang Yang, Hongfei Lin arxiv

Enhancing the moral alignment of Large Language Models (LLMs) is a critical challenge in AI safety. Current alignment techniques often act as superficial guardrails, leaving the intrinsic moral representations of LLMs la…

The Greatest Good Benchmark: Measuring LLMs' Alignment with Utilitarian Moral Dilemmas

2025-03-25 · Giovanni Franco Gabriel Marraffini, Andrés Cotton, Noe Fabian Hsueh, Axel Fridman 외

The question of how to make decisions that maximise the well-being of all persons is very relevant to design language models that are beneficial to humanity and free from harm. We introduce the Greatest Good Benchmark to…

Bounded Morality: Defining the Space of Moral Computation

2026-04-01 · Max Kanwal, Caryn Tran, Patrick Mineault arxiv

Moral cognition has traditionally been modeled as adherence to fixed ethical theories--deontology, consequentialism, virtue ethics--implemented as static rules or value functions. We propose Bounded Morality, a formal fr…

Self-Explaining Hate Speech Detection with Moral Rationales

2026-01-07 · Francielle Vargas, Jackson Trager, Diego Alves, Surendrabikram Thapa 외 arxiv

Hate speech detection models rely on surface-level lexical features, increasing vulnerability to spurious correlations and limiting robustness, cultural contextualization, and interpretability. We propose Supervised Mora…

Hate Speech Detection