paper-with-me

Papers

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs

2026-05-22 · Qitao Tan, Xiaoying Song, Arman Akbari, Arash Akbari, Yanzhi Wang, Xiaoming Zhai, Lingzi Hong, Zhen Xiang, Jin Lu, Geng Yuan arxiv

Current safety alignment of foundation models largely follows a \emph{one-size-fits-all} paradigm, applying the same refusal policy across users and contexts. As a result, models may refuse requests that are unsafe for general users but legitimate for authorized professionals, limiting helpfulness in specialized professional settings. Existing approaches either require costly realignment or rely on inference-time steering that suffers from imprecise control and added latency. To this end, we propose \textsc{Palette}, a modular, controllable, and efficient framework that selectively relaxes refusal behavior on authorized target domains while preserving standard safety elsewhere. Our method identifies a refusal direction via multi-objective search and internalizes it into the model through lightweight adaptation. \textsc{Palette} further supports modular composition: it learns domain-specific safety controls independently and composes them through parameter merging, enabling on-demand multi-domain authorization without retraining. Experiments across four safety benchmarks, multiple model variants, and both LLMs and VLMs show that \textsc{Palette} delivers precise safety control without sacrificing general utility, offering a practical path toward foundation models that adapt to diverse professional needs.

📄 PDF Abstract BibTeX arXiv:2605.24154

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Decoupled Illumination Priors for Spatially Controllable Multi-View Indoor Scene Relighting

2026-07-09 · Chenjian Gao, Linning Xu, Tianfan Xue arxiv

Indoor scene relighting demands photorealism, precise spatial control, and strict multi-view consistency. While diffusion-based image editing models enable semantic lighting manipulation via text prompts, enforcing exact…

Image Editing

Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs

2025-12-19 · Rujiao Long, Yang Li, Xingyao Zhang, Weixun Wang 외 arxiv

Exploration capacity shapes both inference-time performance and reinforcement learning (RL) training for large (vision-) language models, as stochastic sampling often yields redundant reasoning paths with little high-lev…

Reinforcement Learning

Voxify3D: Pixel Art Meets Volumetric Rendering

2025-12-08 · Yi-Chuan Huang, Jiewen Chan, Hao-Jen Chien, Yu-Lun Liu arxiv

Voxel art is a distinctive stylization widely used in games and digital media, yet automated generation from 3D meshes remains challenging due to conflicting requirements of geometric abstraction, semantic preservation, …

Authorize-on-Demand: Dynamic Authorization with Legality-Aware Intellectual Property Protection for VLMs

2026-03-05 · Lianyu Wang, Meng Wang, Huazhu Fu, Daoqiang Zhang arxiv

The rapid adoption of vision-language models (VLMs) has heightened the demand for robust intellectual property (IP) protection of these high-value pretrained models. Effective IP protection should proactively confine mod…

Flexible Portrait Image Editing with Fine-Grained Control

2022-04-04 · Linlin Liu, QiAn Fu, Fei Hou, Ying He

We develop a new method for portrait image editing, which supports fine-grained editing of geometries, colors, lights and shadows using a single neural network model. We adopt a novel asymmetric conditional GAN architect…

Image GenerationSketch-to-Image Translation