paper-with-me

Papers

PopAlign: Diversifying Contrasting Patterns for a More Comprehensive Alignment

2024-10-17 · Zekun Moore Wang, Shawn Wang, Kang Zhu, Jiaheng Liu, Ke Xu, Jie Fu, Wangchunshu Zhou, Wenhao Huang

Alignment of large language models (LLMs) involves training models on preference-contrastive output pairs to adjust their responses according to human preferences. To obtain such contrastive pairs, traditional methods like RLHF and RLAIF rely on limited contrasting patterns, such as varying model variants or decoding temperatures. This singularity leads to two issues: (1) alignment is not comprehensive; and thereby (2) models are susceptible to jailbreaking attacks. To address these issues, we investigate how to construct more comprehensive and diversified contrasting patterns to enhance preference data (RQ1) and verify the impact of the diversification of contrasting patterns on model alignment (RQ2). For RQ1, we propose PopAlign, a framework that integrates diversified contrasting patterns across the prompt, model, and pipeline levels, introducing six contrasting strategies that do not require additional feedback labeling procedures. Regarding RQ2, we conduct thorough experiments demonstrating that PopAlign significantly outperforms existing methods, leading to more comprehensive alignment.

📄 PDF Abstract BibTeX arXiv:2410.13785

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

RLAIF 설명 없음

Similar Papers 제목 키워드 기반

PopAlign: Population-Level Alignment for Fair Text-to-Image Generation

2024-06-28 · Shufan Li, Harkanwar Singh, Aditya Grover

Text-to-image (T2I) models achieve high-fidelity generation through extensive training on large datasets. However, these models may unintentionally pick up undesirable biases of their training data, such as over-represen…

Image GenerationText to Image GenerationText-to-Image Generation

Mining Contrasting Quasi-Clique Patterns

2018-10-03 · Roberto Alonso, Stephan Günnemann

Mining dense quasi-cliques is a well-known clustering task with applications ranging from social networks over collaboration graphs to document analysis. Recent work has extended this task to multiple graphs; i.e. the go…

Clustering

DPP-TTS: Diversifying prosodic features of speech via determinantal point processes

2023-10-23 · Seongho Joo, Hyukhun Koh, Kyomin Jung

With the rapid advancement in deep generative models, recent neural Text-To-Speech(TTS) models have succeeded in synthesizing human-like speech. There have been some efforts to generate speech with various prosody beyond…

DiversityPoint Processestext-to-speechText to Speech

Graph-Aware Contrasting for Multivariate Time-Series Classification

2023-09-11 · Yucheng Wang, Yuecong Xu, Jianfei Yang, Min Wu 외

Contrastive learning, as a self-supervised learning paradigm, becomes popular for Multivariate Time-Series (MTS) classification. It ensures the consistency across different views of unlabeled samples and then learns effe…

ClassificationContrastive LearningSelf-Supervised LearningTime Series+1

Diversifying Question Generation over Knowledge Base via External Natural Questions

2023-09-23 · Shasha Guo, Jing Zhang, Xirui Ke, Cuiping Li 외

Previous methods on knowledge base question generation (KBQG) primarily focus on enhancing the quality of a single generated question. Recognizing the remarkable paraphrasing ability of humans, we contend that diverse te…

DiversityNatural QuestionsQuestion AnsweringQuestion Generation+1