paper-with-me

Papers

SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization

2025-11-09 · Yue Huang, Xiangqi Wang, Xiangliang Zhang arxiv

In high-stakes scenarios-such as self-harm, legal, or medical queries-LLMs must be both trustworthy and helpful. However, these goals often conflict. We propose priority alignment, a new alignment paradigm that enforces a strict "trustworthy-before-helpful" ordering: optimization of helpfulness is conditioned on first meeting trustworthy thresholds (e.g., harmlessness or honesty). To realize this, we introduce Self-Priority Alignment (SPA)-a fully unsupervised framework that generates diverse responses, self-evaluates them and refines them by the model itself, and applies dual-criterion denoising to remove inconsistency and control variance. From this, SPA constructs lexicographically ordered preference pairs and fine-tunes the model using an uncertainty-weighted alignment loss that emphasizes high-confidence, high-gap decisions. Experiments across multiple benchmarks show that SPA improves helpfulness without compromising safety, outperforming strong baselines while preserving general capabilities. Our results demonstrate that SPA provides a scalable and interpretable alignment strategy for critical LLM applications.

📄 PDF Abstract BibTeX arXiv:2511.06222

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Supervised Fine-Tuning Needs to Unlock the Potential of Token Priority

2026-02-01 · Zhanming Shen, Zeyu Qin, Jiaqi Hu, Wentao Ye 외 arxiv

The transition from fitting empirical data to achieving true human utility is fundamentally constrained by a granularity mismatch, where fine-grained autoregressive generation is often supervised by coarse or uniform sig…

DiSCTT: Consensus-Guided Self-Curriculum for Efficient Test-Time Adaptation in Reasoning

2026-03-05 · Mohammad Mahdi Moradi, Sudhir Mudur arxiv

Test-time adaptation offers a promising avenue for improving reasoning performance in large language models without additional supervision, but existing approaches often apply a uniform optimization objective across all …

Reinforcement LearningTest-time Adaptation

Deriving Consensus for Multi-Parallel Corpora: an English Bible Study

2017-11-01 · IJCNLP 2017 11 · Patrick Xia, David Yarowsky

What can you do with multiple noisy versions of the same text? We present a method which generates a single consensus between multi-parallel corpora. By maximizing a function of linguistic features between word pairs, we…

Machine Translation

Improving Interoperability among Defence and National Security Ontologies: Analysis and Evaluation Tasks

2026-08-06 · Jonathon Dilworth, Pedro Giesteira Cotovio, David Herron, Paul Cripps 외 arxiv

The use of ontologies and knowledge graphs is becoming increasingly widespread in the defence and national security domain. Numerous ontologies have been developed through initiatives led by academia, industry, and gover…

Knowledge Graphs

Micro-Structures Graph-Based Point Cloud Registration for Balancing Efficiency and Accuracy

2024-10-29 · Rongling Zhang, Li Yan, Pengcheng Wei, Hong Xie 외

Point Cloud Registration (PCR) is a fundamental and significant issue in photogrammetry and remote sensing, aiming to seek the optimal rigid transformation between sets of points. Achieving efficient and precise PCR pose…

Point Cloud Registration