Residual Stream Analysis of Overfitting And Structural Disruptions
Ensuring that large language models (LLMs) remain both helpful and harmless poses a significant challenge: fine-tuning on repetitive safety datasets, where unsafe prompts are paired with standard refusal templates, often leads to false refusals, in which benign queries are declined. We first quantify this effect, showing that safety data exhibits substantially lower token entropy and 2-gram diversity (0.048) compared to general instruction data. To uncover the root cause, we introduce FlowLens, a stable PCA-based tool for residual-stream geometry analysis, and reveal that higher proportions of safety examples concentrate variance along a few components, reducing representational smoothness and driving false refusals (false refusal rate rises from 63 percent to 84 percent as safety data increases from 0 percent to 40 percent). Guided by these insights, we propose Variance Concentration Loss (VCL), an auxiliary regularizer that penalizes excessive variance concentration in mid-layer residuals. Empirical results demonstrate that VCL reduces false refusals by over 35 percentage points while maintaining or improving performance on general benchmarks such as MMLU and GSM8K.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Statistical vs. Deep Learning Models for Estimating Substance Overdose Excess Mortality in the US
Substance overdose mortality in the United States claimed over 80,000 lives in 2023, with the COVID-19 pandemic exacerbating existing trends through healthcare disruptions and behavioral changes. Estimating excess mortal…
Econometric Analysis of Pandemic Disruption and Recovery Trajectory in the U.S. Rail Freight Industry
To measure the impacts on U.S. rail and intermodal freight by economic disruptions of the 2007-09 Great Recession and the COVID-19 pandemic, this paper uses time series analysis with the AutoRegressive Integrated Moving …
counterfactualTime Series AnalysisPlasma State Monitoring and Disruption Characterization using Multimodal VAEs
When a plasma disrupts in a tokamak, significant heat and electromagnetic loads are deposited onto the surrounding device components. These forces scale with plasma current and magnetic field strength, making disruptions…
counterfactualDiagnosticStructural Design of Convolutional Neural Networks for Steganalysis
Recent studies have indicated that the architectures of convolutional neural networks (CNNs) tailored for computer vision may not be best suited to image steganalysis. In this letter, we report a CNN architecture that ta…
SteganalysisStructural Residual Learning for Single Image Rain Removal
To alleviate the adverse effect of rain streaks in image processing tasks, CNN-based single image rain removal methods have been recently proposed. However, the performance of these deep learning methods largely relies o…
Rain Removal