paper-with-me

홈 › Papers

Polarity-Aware Probing for Quantifying Latent Alignment in Language Models

2025-11-21 · Sabrina Sadiekh, Elena Ericheva, Chirag Agarwal arxiv

Advances in unsupervised probes such as Contrast-Consistent Search (CCS), which reveal latent beliefs without relying on token outputs, raise the question of whether these methods can reliably assess model alignment. We investigate this by examining the sensitivity of CCS to harmful vs. safe statements and by introducing Polarity-Aware CCS (PA-CCS), a method for evaluating whether a model's internal representations remain consistent under polarity inversion. We propose two alignment-oriented metrics, Polar-Consistency and the Contradiction Index, to quantify the semantic robustness of a model's latent knowledge. To validate PA-CCS, we curate two main datasets and one control dataset containing matched harmful-safe sentence pairs constructed using different methodologies (concurrent and antagonistic statements). We apply PA-CCS to 16 language models. Our results show that PA-CCS identifies both architectural and layer-specific differences in the encoding of latent harmful knowledge. Notably, replacing the negation token with a meaningless marker degrades PA-CCS scores for models with well-aligned internal representations, while models lacking robust internal calibration do not exhibit this degradation. Our findings highlight the potential of unsupervised probing for alignment evaluation and emphasize the need to incorporate structural robustness checks into interpretability benchmarks. Code and datasets are available at: https://github.com/SadSabrina/polarity-probing. WARNING: This paper contains potentially sensitive, harmful, and offensive content.

📄 PDF Abstract BibTeX arXiv:2511.21737

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On the Transformation of Latent Space in Fine-Tuned NLP Models

2022-10-23 · Nadir Durrani, Hassan Sajjad, Fahim Dalvi, Firoj Alam

We study the evolution of latent space in fine-tuned NLP models. Different from the commonly used probing-framework, we opt for an unsupervised method to analyze representations. More specifically, we discover latent con…

Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning

2025-10-14 · Noor Islam S. Mohammad, Uluğ Bayazıt arxiv

We treat the internals of generative models as mechanistic objects rather than black boxes. We introduce \textbf{Attribution Graphs} (AGs), which extend GradCAM++ to circuit-level representations, and \textbf{Causal Prob…

Adversarial Robustness

Defending Large Language Models Against Jailbreak Attacks via In-Decoding Safety-Awareness Probing

2026-01-15 · Yinzhi Zhao, Ming Wang, Shi Feng, Xiaocui Yang 외 arxiv

Large language models (LLMs) have achieved impressive performance across natural language tasks and are increasingly deployed in real-world applications. Despite extensive safety alignment efforts, recent studies show th…

Linear Control of Test Awareness Reveals Differential Compliance in Reasoning Models

2025-05-20 · Sahar Abdelnabi, Ahmed Salem

Reasoning-focused large language models (LLMs) sometimes alter their behavior when they detect that they are being evaluated, an effect analogous to the Hawthorne phenomenon, which can lead them to optimize for test-pass…

Safety Alignment

SDSC:A Structure-Aware Metric for Semantic Signal Representation Learning

2025-07-19 · Jeyoung Lee, Hochul Kang arxiv

We propose the Signal Dice Similarity Coefficient (SDSC), a structure-aware metric function for time series self-supervised representation learning. Most Self-Supervised Learning (SSL) methods for signals commonly adopt …

Self-Supervised LearningRepresentation Learning