paper-with-me

홈 › Papers

C$^{2}$R: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse Autoencoders

2026-06-29 · Haoran Jin, Xiting Wang, Shijie Ren, Hong Xie, Defu Lian arxiv

Sparse Autoencoders (SAEs) are widely used to interpret large language models by decomposing activations into sparse, human-understandable features, but scaling to large dictionaries exposes fundamental challenges. Systematic studies reveal pervasive feature splitting that fragments coherent concepts into non-atomic latents and widespread feature absorption that creates arbitrary exceptions in general features, severely compromising latent reliability. These issues stem from inconsistent latent assignment across samples: without cross-sample constraints, per-sample optimization often allows a single underlying concept to be inconsistently distributed across multiple redundant or interfering latents. To address this, we introduce C$^2$R (\underline{\textbf{C}}ross-sample \underline{\textbf{C}}onsistency \underline{\textbf{R}}egularization). C$^2$R explicitly encourages that each semantic feature is consistently represented by a unified latent across the batch by penalizing the co-activation of directionally similar latents. Comprehensive evaluation demonstrates that C$^2$R effectively mitigates both splitting and absorption while, crucially, preserving reconstruction fidelity, providing a principled solution that enhances latent interpretability without degrading model performance. Source code is available at https://github.com/hr-jin/Cross-sample-Consistency-Regularization.

📄 PDF Abstract BibTeX arXiv:2606.30609

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FreqDebias: Towards Generalizable Deepfake Detection via Consistency-Driven Frequency Debiasing

2025-09-26 · Hossein Kashiani, Niloufar Alipour Talemi, Fatemeh Afghah arxiv

Deepfake detectors often struggle to generalize to novel forgery types due to biases learned from limited training data. In this paper, we identify a new type of model bias in the frequency domain, termed spectral bias, …

Representation LearningDomain GeneralizationDeepFake Detection

FreqDebias: Towards Generalizable Deepfake Detection via Consistency-Driven Frequency Debiasing

2025-01-01 · CVPR 2025 1 · Hossein Kashiani, Niloufar Alipour Talemi, Fatemeh Afghah

Deepfake detectors often struggle to generalize to novel forgery types due to biases learned from limited training data. In this paper, we identify a new type of model bias in the frequency domain, termed spectral bi…

DeepFake DetectionDomain GeneralizationFace SwappingRepresentation Learning

Adaptive Dropout: Unleashing Dropout across Layers for Generalizable Image Super-Resolution

2025-01-01 · CVPR 2025 1 · Hang Xu, Jie Huang, Wei Yu, Jiangtong Tan 외

Blind Super-Resolution(blind SR) aims to enhance the model's generalization ability with unknown degradation, yet it still encounters severe overfitting issues. Some previous methods inspired by dropout, which enhanc…

AttributeBlind Super-ResolutionImage RestorationImage Super-Resolution+1

Contrastive Regularization for Semi-Supervised Learning

2022-01-17 · Doyup Lee, Sungwoong Kim, Ildoo Kim, Yeongjae Cheon 외

Consistency regularization on label predictions becomes a fundamental technique in semi-supervised learning, but it still requires a large number of training iterations for high performance. In this study, we analyze tha…

Semi-Supervised Image Classification

Externally Validated Multi-Task Learning via Consistency Regularization Using Differentiable BI-RADS Features for Breast Ultrasound Tumor Segmentation

2025-11-20 · Jingru Zhang, Saed Moradi, Ashirbani Saha arxiv

Multi-task learning can suffer from destructive task interference, where jointly trained models underperform single-task baselines and limit generalization. To improve generalization performance in breast ultrasound-base…

Multi-Task LearningTumor Segmentation