paper-with-me

홈 › Papers

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security

2025-09-08 · Yanrui Du, Fenglei Fan, Sendong Zhao, Jiawei Cao, Ting Liu, Bing Qin arxiv

As Large Language Models (LLMs) increasingly permeate human life, their security has emerged as a critical concern, particularly their ability to maintain harmless responses to malicious instructions. Although extensive methods have improved LLMs' security, they often lead to conservative, rejection-oriented responses that compromise practical usability. This presents a key challenge: how to advance the Pareto frontier between LLMs' usability and security, rather than necessitate a trade-off between them. To address this, we propose the MoGU framework, in which the intra-layer router dynamically allocates weights by sensing hidden states, thereby balancing the contributions of security-optimized and usability-optimized variants. Despite its initial potential, the MoGU framework faces limitations such as parameter redundancy and performance bottlenecks. To overcome these, we further propose an improved MoGU_v2 framework that establishes a tighter coupling between the routers and hidden states. In MoGU_v2, routers are embedded only in layers encoding highly classifiable security features, and backbone modules are activated during router optimization to enable bidirectional adaptation. MoGU_V2 exhibits strong adaptability and stable improvements across various series of LLMs, including mainstream LLMs serving as brains in various applications, on-device LLMs optimized for resource-constrained scenarios, and reasoning LLMs tailored for user interpretability. Meanwhile, even facing risks introduced by Instruction Fine-tuning, MoGU_v2 can easily restore security without compromising the task performance gains via a simple data-mix strategy. These comprehensive improvements highlight MoGU_V2 as a robust and versatile solution for mitigating security risks in real-world applications.

📄 PDF Abstract BibTeX arXiv:2509.06807

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MoGU: A Framework for Enhancing Safety of Open-Sourced LLMs While Preserving Their Usability

2024-05-23 · Yanrui Du, Sendong Zhao, Danyang Zhao, Ming Ma 외

Large Language Models (LLMs) are increasingly deployed in various applications. As their usage grows, concerns regarding their safety are rising, especially in maintaining harmless responses when faced with malicious ins…

Towards Higher Pareto Frontier in Multilingual Machine Translation

2023-05-25 · Yichong Huang, Xiaocheng Feng, Xinwei Geng, Baohang Li 외

Multilingual neural machine translation has witnessed remarkable progress in recent years. However, the long-tailed distribution of multilingual corpora poses a challenge of Pareto optimization, i.e., optimizing for some…

Knowledge DistillationMachine TranslationTranslation

Towards Efficient NLP: A Standard Evaluation and A Strong Baseline

2021-10-13 · NAACL 2022 7 · Xiangyang Liu, Tianxiang Sun, Junliang He, Jiawen Wu 외

Supersized pre-trained language models have pushed the accuracy of various natural language processing (NLP) tasks to a new state-of-the-art (SOTA). Rather than pursuing the reachless SOTA accuracy, more and more researc…

Approximating Pareto Frontiers in Stochastic Multi-Objective Optimization via Hashing and Randomization

2026-04-01 · Jinzhao Li, Nan Jiang, Yexiang Xue arxiv

Stochastic Multi-Objective Optimization (SMOO) is critical for decision-making trading off multiple potentially conflicting objectives in uncertain environments. SMOO aims at identifying the Pareto frontier, which contai…

Ranking by Momentum based on Pareto ordering of entities

2021-11-25 · Tomasz Imielinski

Given a set of changing entities, which ones are the most uptrending over some time T? Which entities are standing out as the biggest movers? To answer this question we define the concept of momentum. Two parameters - ab…