paper-with-me

Papers

Unintended Impacts of LLM Alignment on Global Representation

2024-02-22 · Michael J. Ryan, William Held, Diyi Yang

Before being deployed for user-facing applications, developers align Large Language Models (LLMs) to user preferences through a variety of procedures, such as Reinforcement Learning From Human Feedback (RLHF) and Direct Preference Optimization (DPO). Current evaluations of these procedures focus on benchmarks of instruction following, reasoning, and truthfulness. However, human preferences are not universal, and aligning to specific preference sets may have unintended effects. We explore how alignment impacts performance along three axes of global representation: English dialects, multilingualism, and opinions from and about countries worldwide. Our results show that current alignment procedures create disparities between English dialects and global opinions. We find alignment improves capabilities in several languages. We conclude by discussing design decisions that led to these unintended impacts and recommendations for more equitable preference tuning. We make our code and data publicly available on Github.

📄 PDF Abstract BibTeX arXiv:2402.15018

Code (1)

salt-nlp/unintended-impacts-of-alignment 공식 구현 pytorch

Tasks

Instruction Following

Methods 이 논문이 사용한 방법론

Focus 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Can Probabilistic Feedback Drive User Impacts in Online Platforms?

2024-01-10 · Jessica Dai, Bailey Flanigan, Nika Haghtalab, Meena Jagadeesan 외

A common explanation for negative user impacts of content recommender systems is misalignment between the platform's objective and user welfare. In this work, we show that misalignment in the platform's objective is not …

Recommendation Systems

Language Models Resist Alignment: Evidence From Data Compression

2024-06-10 · Jiaming Ji, Kaile Wang, Tianyi Qiu, Boyuan Chen 외

Large language models (LLMs) may exhibit unintended or undesirable behaviors. Recent works have concentrated on aligning LLMs to mitigate harmful outputs. Despite these efforts, some anomalies indicate that even a well-c…

Data Compression

Self-Regularization with Latent Space Explanations for Controllable LLM-based Classification

2025-02-19 · Xuansheng Wu, Wenhao Yu, Xiaoming Zhai, Ninghao Liu

Modern text classification methods heavily rely on contextual embeddings from large language models (LLMs). Compared to human-engineered features, these embeddings provide automatic and effective representations for clas…

ClassificationFairnesstext-classificationText Classification

NLP-based Cross-Layer 5G Vulnerabilities Detection via Fuzzing Generated Run-Time Profiling

2023-05-14 · Zhuzhu Wang, Ying Wang

The effectiveness and efficiency of 5G software stack vulnerability and unintended behavior detection are essential for 5G assurance, especially for its applications in critical infrastructures. Scalability and automatio…

Demographics Should Not Be the Reason of Toxicity: Mitigating Discrimination in Text Classifications with Instance Weighting

2020-04-29 · ACL 2020 6 · Guanhua Zhang, Bing Bai, Junqi Zhang, Kun Bai 외

With the recent proliferation of the use of text classifications, researchers have found that there are certain unintended biases in text classification datasets. For example, texts containing some demographic identity-t…

Abusive LanguageGeneral ClassificationSelection biastext-classification+1