paper-with-me

홈 › Papers

Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety

2025-09-16 · Denis Janiak, Julia Moska, Dawid Motyka, Karolina Seweryn, Paweł Walkowiak, Bartosz Żuk, Arkadiusz Janz arxiv

Large language models (LLMs) require careful alignment to balance competing objectives - factuality, safety, conciseness, proactivity, and diversity. Existing studies focus on individual techniques or specific dimensions, lacking a holistic assessment of the inherent trade-offs. We propose a unified evaluation framework that compares LLM alignment methods (PPO, DPO, ORPO, KTO) across these five axes, using both in-distribution and out-of-distribution datasets. Leveraging a specialized LLM-as-Judge prompt, validated through human studies, we reveal that DPO and KTO excel in factual accuracy, PPO and DPO lead in safety, and PPO best balances conciseness with proactivity. Our findings provide insights into trade-offs of common alignment methods, guiding the development of more balanced and reliable LLMs.

📄 PDF Abstract BibTeX arXiv:2509.12936

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Rethinking Inverse Reinforcement Learning: from Data Alignment to Task Alignment

2024-10-31 · Weichao Zhou, Wenchao Li

Many imitation learning (IL) algorithms use inverse reinforcement learning (IRL) to infer a reward function that aligns with the demonstration. However, the inferred reward functions often fail to capture the underlying …

Imitation LearningTransfer Learning

Rethinking Cross-Modal Fine-Tuning: Optimizing the Interaction Between Feature Alignment and Target Fitting

2026-01-26 · Trong Khiem Tran, Manh Cuong Dao, Phi Le Nguyen, Thao Nguyen Truong 외 arxiv

Adapting pre-trained models to unseen feature modalities has become increasingly important due to the growing need for cross-disciplinary knowledge integration. A key challenge here is how to align the representation of …

Rethinking One-Step Image Editing through ChordEdit: Reproduction, Simplification, and New Insights

2026-06-12 · Minghan Li, Jeremy Moebel, Mengyu Wang arxiv

One-step image editing is important for making text-guided editing fast, practical, and easy to deploy, but its underlying mechanism is still not fully understood. We revisit ChordEdit through reproduction, ablation, and…

Image Editing

Rethinking Alignment in Video Super-Resolution Transformers

2022-07-18 · Shuwei Shi, Jinjin Gu, Liangbin Xie, Xintao Wang 외

The alignment of adjacent frames is considered an essential operation in video super-resolution (VSR). Advanced VSR models, including the latest VSR Transformers, are generally equipped with well-designed alignment modul…

Common Sense ReasoningSuper-ResolutionVideo Super-Resolution

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding

2025-11-17 · Jiaze Li, Hao Yin, Wenhui Tan, Jingyang Chen 외 arxiv

Self-reflection mechanisms that rely on purely text-based rethinking processes perform well in most multimodal tasks. However, when directly applied to long-form video understanding scenarios, they exhibit clear limitati…

Reinforcement Learning