paper-with-me

홈 › Papers

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms

2026-05-08 · Linh Le, David Williams-King, Mohamed Amine Merzouk, Aton Kamanda, Adam Oberman arxiv

Current adversarial robustness methods for large language models require extensive datasets of harmful prompts (thousands to hundreds of thousands of examples), yet remain vulnerable to novel attack vectors and distributional shifts. We propose Latent Personality Alignment (LPA), a sample-efficient defense that achieves robustness by training models on abstract personality traits rather than specific harmful behaviors. Using fewer than 100 trait statements and latent adversarial training, LPA achieves comparable attack success rates to methods trained on 150k+ examples, while maintaining superior utility. Critically, LPA generalizes better to unseen attack distributions, reducing misclassification rates by 2.6x compared to baseline across six harm benchmarks -- without ever seeing harmful examples during training. Our results demonstrate that personality-based alignment offers a principled approach to building robust defenses with minimal cost.

📄 PDF Abstract BibTeX arXiv:2605.08496

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial Robustness

Similar Papers 제목 키워드 기반

Large-Language-Models-as-a-Judge in Theory-Agnostic Adaptive Metric-Alignment for Prototypical Networks in Personality Recognition

2026-07-09 · Jing Jie Tan, Ban-Hoe Kwan, Danny Wee-Kiat Ng, Yan-Chai Hum 외 arxiv

Personality recognition has traditionally been constrained by theory-dependent formulations, where models are trained to fit predefined psychological taxonomies rather than uncovering shared underlying behavioral structu…

Metric Learning

Efficient Safety Alignment of Language Models via Latent Personality Traits

2026-07-08 · Mohamed Amine Merzouk, Nolan Smyth, Damiano Fornasiere, Linh Le 외 arxiv

Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives. Latent Adversarial Training (LAT) is among the most effective defenses, bu…

The Geometry of Persona: Disentangling Personality from Reasoning in Large Language Models

2025-12-08 · Zhixiang Wang arxiv

Background: The deployment of personalized Large Language Models (LLMs) is currently constrained by the stability-plasticity dilemma. Prevailing alignment methods, such as Supervised Fine-Tuning (SFT), rely on stochastic…

Rediscovering the Latent Dimensions of Personality with Large Language Models as Trait Descriptors

2024-09-16 · Joseph Suh, Suhong Moon, Minwoo Kang, David M. Chan

Assessing personality traits using large language models (LLMs) has emerged as an interesting and challenging area of research. While previous methods employ explicit questionnaires, often derived from the Big Five model…

Descriptive

Dynamic Personality Adaptation in Large Language Models via State Machines

2026-02-25 · Leon Pielage, Ole Hätscher, Mitja Back, Bernhard Marschall 외 arxiv

The inability of Large Language Models (LLMs) to modulate their personality expression in response to evolving dialogue dynamics hinders their performance in complex, interactive contexts. We propose a model-agnostic fra…