paper-with-me

Papers

LMSM: LLM Security Framework Inspired by Linux Security Modules

2026-08-26 · XiuYu Zhang, Bonan Ruan, Junfeng Fang, An Zhang, Tat-Seng Chua, Zhenkai Liang hf

Large language models (LLMs) are increasingly deployed with layered defenses, yet malicious prompts can still bypass them. Interpretability methods can expose model-internal signals along the generation path that could inform enforcement, but these signals are not security controls by themselves. Deployments that adapt them for safety typically couple each signal to its own calibration, policy logic, and intervention code, so each new artifact creates integration work instead of strengthening a shared defense. We present Language Model Security Modules (LMSM), a security framework that adapts the separation behind Linux Security Modules (LSM) to LLM serving. In LMSM, a selected security backend exposes calibrated evidence, a versioned policy evaluates active rules over trusted per-request context, and a separate gate authorizes buffered output release. This design separates mediation correctness from policy effectiveness, and it allows backend, rule, or schedule changes without rebuilding request handling or enforcement. Our prototype shows the separation working in practice: with Hugging Face Transformers and continuously batched vLLM, the same substrate hosts artifact-backed sparse autoencoder (SAE) and transcoder deployments and task-fitted dense probes, preserves request-specific decisions under scheduler churn, and selectively enforces and composes multiple rules per request. On Qwen3-4B, LMSM-Checkpoint reduces HarmBench attack success rate from 39.20% to 3.32%, with XSTest false refusals rising from 2.40% to 4.40%, while retaining 98.14% of the throughput of a matched serving path that performs no monitoring work at 32 active sequences. LMSM gives advances in interpretability and model-internal analysis a common path to runtime enforcement.

📄 PDF Abstract BibTeX arXiv:2608.25697

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Security Hardening Using FABRIC: Implementing a Unified Compliance Aggregator for Linux Servers

2026-01-01 · Sheldon Paul, Izzat Alsmadi arxiv

This paper presents a unified framework for evaluating Linux security hardening on the FABRIC testbed through aggregation of heterogeneous security auditing tools. We deploy three Ubuntu 22.04 nodes configured at baselin…

Machine Learning-Based Security Policy Analysis

2024-12-30 · Krish Jain, Joann Sum, Pranav Kapoor, Amir Eaman

Security-Enhanced Linux (SELinux) is a robust security mechanism that enforces mandatory access controls (MAC), but its policy language's complexity creates challenges for policy analysis and management. This research in…

Anomaly Detection

Towards Reliable Local Security Agents: Verifiable Post-Training for Linux Privilege Escalation

2026-03-18 · Philipp Normann, Andreas Happe, Jürgen Cito, Daniel Arp arxiv

LLM agents are becoming increasingly important in the security domain, but leading systems are often closed-source, cloud-based, hard to reproduce or use with sensitive code. This creates a need for small, local models t…

Reinforcement Learning

LLM in the Shell: Generative Honeypots

2023-08-31 · Muris Sladić, Veronica Valeros, Carlos Catania, Sebastian Garcia

Honeypots are essential tools in cybersecurity for early detection, threat intelligence gathering, and analysis of attacker's behavior. However, most of them lack the required realism to engage and fool human attackers l…

What Do They Fix? LLM-Aided Categorization of Security Patches for Critical Memory Bugs

2025-09-26 · Xingyu Li, Juefei Pu, Yifan Wu, Xiaochen Zou 외 arxiv

Open-source software projects are foundational to modern software ecosystems, with the Linux kernel standing out as a critical exemplar due to its ubiquity and complexity. Although security patches are continuously integ…