paper-with-me

홈 › Papers

Efficient Safety Alignment of Language Models via Latent Personality Traits

2026-07-08 · Mohamed Amine Merzouk, Nolan Smyth, Damiano Fornasiere, Linh Le, David Williams-King, Adam Oberman arxiv

Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives. Latent Adversarial Training (LAT) is among the most effective defenses, but can degrade utility and requires training on large datasets of harmful prompts. We introduce Latent Personality Alignment (LPA), which replaces explicit harm refusal with adversarial training on just 66 harm-agnostic statements drawn from psychometric personality literature. We hypothesize that personality-anchored representations share latent structure with harm avoidance, so adversarially stabilizing them implicitly constrains the subspace exploited by jailbreak attacks. LPA achieves near-zero attack success rates on HarmBench across direct requests and five jailbreak methods, despite never seeing harmful content during training and no loss of performance on standard benchmarks. Moreover, the training process is lightweight; the entire procedure completes in minutes on a single GPU and uses 75x fewer examples than standard LAT. Extensive ablations demonstrate the robustness, efficiency, and generalization of our method.

📄 PDF Abstract BibTeX arXiv:2607.07918

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Better Angels of Machine Personality: How Personality Relates to LLM Safety

2024-07-17 · Jie Zhang, Dongrui Liu, Chen Qian, Ziyue Gan 외

Personality psychologists have analyzed the relationship between personality and safety behaviors in human society. Although Large Language Models (LLMs) demonstrate personality traits, the relationship between personali…

FairnessSafety Alignment

Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models

2025-09-19 · Stephen Fitz, Peter Romero, Steven Basart, Sipeng Chen 외 arxiv

Large Language Models increasingly mediate high-stakes interactions, intensifying research on their capabilities and safety. While recent work has shown that LLMs exhibit consistent and measurable synthetic personality t…

AI-exhibited Personality Traits Can Shape Human Self-concept through Conversations

2026-01-19 · Jingshu Li, Tianqi Song, Nattapat Boonprakong, Zicheng Zhu 외 arxiv

Recent Large Language Model (LLM) based AI can exhibit recognizable and measurable personality traits during conversations to improve user experience. However, as human understandings of their personality traits can be a…

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms

2026-05-08 · Linh Le, David Williams-King, Mohamed Amine Merzouk, Aton Kamanda 외 arxiv

Current adversarial robustness methods for large language models require extensive datasets of harmful prompts (thousands to hundreds of thousands of examples), yet remain vulnerable to novel attack vectors and distribut…

Adversarial Robustness

Do GPT Language Models Suffer From Split Personality Disorder? The Advent Of Substrate-Free Psychometrics

2024-08-14 · Peter Romero, Stephen Fitz, Teruo Nakatsuma

Previous research on emergence in large language models shows these display apparent human-like abilities and psychological latent traits. However, results are partly contradicting in expression and magnitude of these la…

Language ModelingLanguage Modelling