paper-with-me

홈 › Papers

Optimizing Safe and Aligned Language Generation: A Multi-Objective GRPO Approach

2025-03-26 · Xuying Li, Zhuo Li, Yuji Kosuga, Victor Bian

Aligning large language models (LLMs) with human values and safety constraints is challenging, especially when objectives like helpfulness, truthfulness, and avoidance of harm conflict. Reinforcement Learning from Human Feedback (RLHF) has achieved notable success in steering models, but is complex and can be unstable. Recent approaches such as Direct Preference Optimization (DPO) simplify preference-based fine-tuning but may introduce bias or trade-off certain objectives~\cite{dpo}. In this work, we propose a Group Relative Policy Optimization (GRPO) framework with a multi-label reward regression model to achieve safe and aligned language generation. The GRPO algorithm optimizes a policy by comparing groups of sampled responses, eliminating the need for a separate value critic and improving training efficiency~\cite{grpo}. We train a reward model to predict multiple alignment scores (e.g., safety, helpfulness, etc.), which are combined into a single reward signal. We provide a theoretical derivation for using this learned multi-aspect reward within GRPO and discuss its advantages and limitations. Empirically, our approach improves all the safety and quality metrics evaluated in language generation tasks on model scales (0.5B, 7B, and 14B parameters), demonstrating a robust balance of objectives. We compare GRPO to PPO-based RLHF and DPO, highlighting that GRPO achieves alignment with significantly lower computational cost and explicit multi-objective handling. \textbf{We will open-source all trained models at https://huggingface.co/hydroxai.

📄 PDF Abstract BibTeX arXiv:2503.21819

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Methods 이 논문이 사용한 방법론

DPO 설명 없음

Similar Papers 제목 키워드 기반

ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models

2025-12-06 · Somnath Banerjee, Sayan Layek, Sayantan Adak, Mykola Pechenizkiy 외 arxiv

Current language model safety paradigms often fall short in emotionally charged or high-stakes settings, where refusal-only approaches may alienate users and naive compliance can amplify risk. We propose ProSocialAlign, …

State-Dependent Safety Failures in Multi-Turn Language Model Interaction

2026-03-15 · Pengcheng Li, Jie Zhang, Tianwei Zhang, Han Qiu 외 arxiv

Safety alignment in large language models is typically evaluated under isolated queries, yet real-world use is inherently multi-turn. Although multi-turn jailbreaks are empirically effective, the structure of conversatio…

Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching

2024-05-22 · Weixiang Zhao, Yulin Hu, Zhuojun Li, Yang Deng 외

Safety alignment of large language models (LLMs) has been gaining increasing attention. However, current safety-aligned LLMs suffer from the fragile and imbalanced safety mechanisms, which can still be induced to generat…

Safety Alignment

Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation

2025-11-12 · Xin Zhao, Xiaojun Chen, Bingshan Liu, Zeyao Liu 외 arxiv

Generative vision-language models like Stable Diffusion demonstrate remarkable capabilities in creative media synthesis, but they also pose substantial risks of producing unsafe, offensive, or culturally inappropriate co…

Text-to-Image Generation

Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF

2026-04-20 · Yuan Fang, Yiming Luo, Aimin Zhou, Fei Tan arxiv

Ensuring the safety of large language models (LLMs) requires robust red teaming, yet the systematic synthesis of high-quality toxic data remains under-explored. We propose Reverse Constitutional AI (R-CAI), a framework f…

Reinforcement LearningRed Teaming