paper-with-me

Papers

When the Model Said 'No Comment', We Knew Helpfulness Was Dead, Honesty Was Alive, and Safety Was Terrified

2026-02-07 · Gautam Siddharth Kashyap, Mark Dras, Usman Naseem arxiv

Large Language Models (LLMs) need to be in accordance with human values-being helpful, harmless, and honest (HHH)-is important for safe deployment. Existing works use Supervised Fine-Tuning (SFT) and Mixture-of-Experts (MoE) to align LLMs. However, these works face challenges in multi-objective settings, such as SFT leading to interference between conflicting objectives, while MoEs suffer from miscalibrated routing. We term this failure mode Axis Collapse, marked by (1) disjoint feature spaces causing catastrophic forgetting, and (2) unreliable inference from misrouted experts. To resolve this, we propose AlignX, a two-stage framework. Stage 1 uses prompt-injected fine-tuning to extract axis-specific task features, mitigating catastrophic forgetting. Stage 2 deploys a MoCaE module that calibrates expert routing using fractal and natural geometry, improving inference reliability. AlignX achieves significant gains on Alpaca (Helpfulness), BeaverTails (Harmlessness), and TruthfulQA (Honesty), with +171.5% win rate, +110.1% in truthfulness-informativeness, and 4.3% fewer safety violations. It also reduces latency and memory usage by over 35% compared to prior MoEs. Results across four LLMs validate its generalizability.

📄 PDF Abstract BibTeX arXiv:2602.07381

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

How to be Helpful on Online Support Forums?

2022-07-01 · NAACL (WNU) 2022 7 · Zhilin Wang, Pablo E. Torres

Internet forums such as Reddit offer people a platform to ask for advice when they encounter various issues at work, school or in relationships. Telling helpful comments apart from unhelpful comments to these advice-seek…

Modeling and Prediction of Online Product Review Helpfulness: A Survey

2018-07-01 · ACL 2018 7 · Gerardo Ocampo Diaz, Vincent Ng

As the amount of free-form user-generated reviews in e-commerce websites continues to increase, there is an increasing need for automatic mechanisms that sift through the vast amounts of user reviews and identify quality…

Recommendation SystemsSurvey

How to be Helpful on Online Support Forums?

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Internet forums such as Reddit offer people a platform to ask for advice when they encounter various issues at work, school or in relationships. Telling helpful comments apart from unhelpful comments to these advice-seek…

Classification of comment helpfulness to improve knowledge sharing among medical practitioners.

2016-06-01 · WS 2016 6 · Pierre Andr{\'e} M{\'e}nard, Caroline Barri{\`e}re
General ClassificationOpinion Mining

Slavic Forest, Norwegian Wood

2017-04-01 · WS 2017 4 · Rudolf Rosa, Daniel Zeman, David Mare{\v{c}}ek, Zden{\v{e}}k {\v{Z}}abokrtsk{\'y}

We once had a corp, or should we say, it once had us They showed us its tags, isn{'}t it great, unified tags They asked us to parse and they told us to use everything So we looked around and we noticed there was near not…

Dependency ParsingMachine TranslationWord Alignment