paper-with-me

홈 › Papers

Guarding the Meaning: Self-Supervised Training for Semantic Robustness in Guard Models

2025-11-06 · Cristina Pinneri, Christos Louizos arxiv

Guard models are a critical component of LLM safety, but their sensitivity to superficial linguistic variations remains a key vulnerability. We show that even meaning-preserving paraphrases can cause large fluctuations in safety scores, revealing a lack of semantic grounding. To address this, we introduce a practical, self-supervised framework for improving the semantic robustness of guard models. Our method leverages paraphrase sets to enforce prediction consistency using a novel, skew-aware aggregation strategy for robust target computation. Notably, we find that standard aggregation methods like mean and median can degrade safety, underscoring the need for skew-aware alternatives. We analyze six open-source guard models and show that our approach reduces semantic variability across paraphrases by ~58%, improves benchmark accuracy by ~2.5% on average, and generalizes to unseen stylistic variations. Intriguingly, we discover a bidirectional relationship between model calibration and consistency: our robustness training improves calibration by up to 40%, revealing a fundamental connection between these properties. These results highlight the value of treating semantic consistency as a first-class training objective and provide a scalable recipe for building more reliable guard models.

📄 PDF Abstract BibTeX arXiv:2511.10665

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Distilling Word Meaning in Context from Pre-trained Language Models

2021-11-01 · Findings (EMNLP) 2021 11 · Yuki Arase, Tomoyuki Kajiwara

In this study, we propose a self-supervised learning method that distils representations of word meaning in context from a pre-trained masked language model. Word representations are the basis for context-aware lexical s…

Language ModelingLanguage ModellingSelf-Supervised LearningSemantic Textual Similarity+2

SMCLM: Semantically Meaningful Causal Language Modeling for Autoregressive Paraphrase Generation

2025-07-04 · Michał Perełkiewicz, Sławomir Dadas, Rafał Poświata arxiv

This article introduces semantically meaningful causal language modeling (SMCLM), a selfsupervised method of training autoregressive models to generate semantically equivalent text. Our approach involves using semantical…

Paraphrase Generation

Multi-Modal Self-Supervised Semantic Communication

2025-03-18 · Hang Zhao, Hongru Li, Dongfang Xu, Shenghui Song 외

Semantic communication is emerging as a promising paradigm that focuses on the extraction and transmission of semantic meanings using deep learning techniques. While current research primarily addresses the reduction of …

Self-Supervised LearningSemantic Communication

Self-Supervised Learning of Pretext-Invariant Representations

2019-12-04 · CVPR 2020 6 · Ishan Misra, Laurens van der Maaten

The goal of self-supervised learning from images is to construct image representations that are semantically meaningful via pretext tasks that do not require semantic annotations for a large training set of images. Many …

Contrastive Learningobject-detectionObject DetectionRepresentation Learning+3

Implicit Surface Contrastive Clustering for LiDAR Point Clouds

2023-01-01 · CVPR 2023 1 · Zaiwei Zhang, Min Bai, Erran Li

Self-supervised pretraining on large unlabeled datasets has shown tremendous success on improving the task performance of many computer vision tasks. However, such techniques have not been widely used for outdoor LiD…

3D Object DetectionClusteringContrastive Learningobject-detection+4