paper-with-me

Papers

The Geometry of Harmfulness in LLMs through Subconcept Probing

2025-07-23 · McNair Shah, Saleena Angeline, Adhitya Rajendra Kumar, Naitik Chheda, Kevin Zhu, Vasu Sharma, Sean O'Brien, Will Cai arxiv

Recent advances in large language models (LLMs) have intensified the need to understand and reliably curb their harmful behaviours. We introduce a multidimensional framework for probing and steering harmful content in model internals. For each of 55 distinct harmfulness subconcepts (e.g., racial hate, employment scams, weapons), we learn a linear probe, yielding 55 interpretable directions in activation space. Collectively, these directions span a harmfulness subspace that we show is strikingly low-rank. We then test ablation of the entire subspace from model internals, as well as steering and ablation in the subspace's dominant direction. We find that dominant direction steering allows for near elimination of harmfulness with a low decrease in utility. Our findings advance the emerging view that concept subspaces provide a scalable lens on LLM behaviour and offer practical tools for the community to audit and harden future generations of language models.

📄 PDF Abstract BibTeX arXiv:2507.21141

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Correcting Performance Estimation Bias in Imbalanced Classification with Minority Subconcepts

2026-04-28 · Taylor Maxson, Roberto Corizzo, Yaning Wu, Nathalie Japkowicz 외 arxiv

Class-level evaluation can conceal substantial performance disparities across subconcepts within the same class, causing models that perform well on average to fail on specific subpopulations. Prior work has shown that c…

False Sense of Security: Why Probing-based Malicious Input Detection Fails to Generalize

2025-09-04 · Cheng Wang, Zeming Wei, Qin Liu, Muhao Chen arxiv

Large Language Models (LLMs) can comply with harmful instructions, raising serious safety concerns despite their impressive capabilities. Recent work has leveraged probing-based approaches to study the separability of ma…

LLMs Encode Harmfulness and Refusal Separately

2025-07-16 · Jiachen Zhao, Jing Huang, Zhengxuan Wu, David Bau 외

LLMs are trained to refuse harmful instructions, but do they truly understand harmfulness beyond just refusing? Prior work has shown that LLMs' refusal behaviors can be mediated by a one-dimensional subspace, i.e., a ref…

AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulness

2025-07-02 · Zixin Chen, Hongzhan Lin, Kaixin Li, Ziyang Luo 외 arxiv

The proliferation of multimodal memes in the social media era demands that multimodal Large Language Models (mLLMs) effectively understand meme harmfulness. Existing benchmarks for assessing mLLMs on harmful meme underst…

Mitigating Fine-tuning Risks in LLMs via Safety-Aware Probing Optimization

2025-05-22 · Chengcan Wu, Zhixin Zhang, Zeming Wei, Yihao Zhang 외

The significant progress of large language models (LLMs) has led to remarkable achievements across numerous applications. However, their ability to generate harmful content has sparked substantial safety concerns. Despit…

Safety Alignment