paper-with-me

홈 › Papers

Intrinsic Structure as a Proxy for Saliency: SVD-Based Weight Preservation for Mixed-Precision Quantization in Large Language Models

2025-12-01 · Shashank Landge, Abhishek Patil, Tejas kamble, Bhushan Buddhivant, Priyanka Joshi arxiv

As Large Language Models (LLMs) continue to scale in parameter count, deploying them on commodity hardware has become increasingly challenging. Post-Training Quantization (PTQ) addresses this by reducing the precision of model weights, typically to 4-bit or lower. However, uniform quantization often leads to significant performance degradation due to the presence of ``outlier features'' -- weights that, while few in number, are critical for maintaining model accuracy. Current state-of-the-art methods such as AWQ (Activation-aware Weight Quantization) and SpQR (Sparse Quantization Representations) rely on calibration data to identify these salient weights via activation magnitudes or Hessian sensitivity. In scenarios where data privacy is paramount or calibration data is unavailable, these methods are inapplicable. In this work, we propose a data-free, structure-aware hypothesis: that the weights identified as Principal Components via Singular Value Decomposition (SVD) are intrinsically important to the model's downstream performance. We introduce a novel selection heuristic that preserves the top-$k$ weights aligned with the principal components in FP32, while aggressively quantizing the residual weights. We compare our method against activation-aware (AWQ) and second-order (SpQR) methods across GLUE benchmarks (MRPC, RTE, QNLI) using a DistilBERT backbone. Our experiments reveal that structural importance is highly correlated with functional importance. On the challenging RTE task, our SVD-based method achieves an accuracy of 66.06\%, outperforming both AWQ (65.34\%) and SpQR (65.34\%) at high protection budgets, validating that intrinsic matrix structure can serve as a robust proxy for weight saliency without the need for forward passes or calibration data.

📄 PDF Abstract BibTeX arXiv:2512.01343

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Preservation Is Not Enough for Width Growth: Regime-Sensitive Selection of Dense LM Warm Starts

2026-04-05 · Eren Unlu arxiv

Width expansion offers a practical route to reuse smaller causal-language-model checkpoints, but selecting a widened warm start is not solved by zero-step preservation alone. We study dense width growth as a candidate-se…

Graph-Theoretic Spatiotemporal Context Modeling for Video Saliency Detection

2017-07-25 · Lina Wei, Fangfang Wang, Xi Li, Fei Wu 외

As an important and challenging problem in computer vision, video saliency detection is typically cast as a spatiotemporal context modeling problem over consecutive frames. As a result, a key issue in video saliency dete…

Saliency DetectionVideo Saliency Detection

Saliency-Motion Guided Trunk-Collateral Network for Unsupervised Video Object Segmentation

2025-04-08 · Xiangyu Zheng, Wanyun Li, Songcheng He, Jianping Fan 외

Recent mainstream unsupervised video object segmentation (UVOS) motion-appearance approaches use either the bi-encoder structure to separately encode motion and appearance features, or the uni-encoder structure for joint…

Optical Flow EstimationSalient Object DetectionSemantic SegmentationUnsupervised Video Object Segmentation+3

Scene Parameter Saliency via Differentiable Light Transport

2026-07-23 · Linas Beresna, Eugene Fiume arxiv

Gradient-based saliency methods reveal which input features most influence a neural network's output, and are a standard tool for model interpretability. We observe that differentiable renderers, which are conventionally…

Scene Understanding

SCENE: Semantic-aware Codec Enhancement with Neural Embeddings

2026-01-29 · Han-Yu Lin, Li-Wei Chen, Hung-Shin Lee arxiv

Compression artifacts from standard video codecs often degrade perceptual quality. We propose a lightweight, semantic-aware pre-processing framework that enhances perceptual fidelity by selectively addressing these disto…