Papers Neural Network Compression
“Neural Network Compression” 태그가 달린 논문 215편 · 필터 해제
M-Fibration Theory with Applications to Weighted Graphs
The purpose of this paper is to provide a general, comprehensive, theoretical framework that allows one to deal with fibrations on graphs labelled on a commutative monoid. This is a genuine extension of the theory of gra…
Neural Network CompressionOn the Applicability of Safety Nets: A Safety-By-Design Solution for Certifying Neural Networks
The integration of Artificial Intelligence (AI) in safety-critical aviation systems presents significant challenges for certification and deployment. Aviation, often regarded as the safest form of transportation, relies …
Neural Network CompressionEvoLP: Self-Evolving Latency Predictor for Model Compression in Real-Time Edge Systems
Edge devices are increasingly utilized for deploying deep learning applications on embedded systems. The real-time nature of many applications and the limited resources of edge devices necessitate latency-targeted neural…
Neural Network CompressionModel CompressionHierarchical Reinforcement Learning for Neural Network Compression (HiReLC): Pruning and Quantization
We present HiReLC, a hierarchical ensemble-reinforcement learning framework for automated joint quantization and structured pruning of deep neural networks. The framework decomposes the compression search across two leve…
Hierarchical Reinforcement LearningNeural Network CompressionActive LearningHybrid Compression: Integrating Pruning and Quantization for Optimized Neural Networks
Deep neural networks have witnessed remarkable advancements in recent years and have become integral to various applications. However, alongside these developments, training and deployment of neural network models on emb…
Neural Network CompressionModel CompressionNeural Network Compression by Approximate Differential Equivalence
Neural network compression is commonly achieved by pruning parameters based on local importance scores, e.g., magnitude-based pruning. We propose a complementary approach that compresses models by aggregating neurons wit…
Neural Network CompressionFast Tensorization of Neural Networks via Slice-wise Feature Distillation
We propose a scalable tensorization framework for neural network compression based on slice-wise feature distillation. Unlike conventional tensor decomposition methods that rely on costly global finetuning, our approach …
Neural Network CompressionRobust Basis Spline Decoupling for the Compression of Transformer Models
Decoupling is a powerful modeling paradigm for representing multivariate functions as compositions of linear transformations and univariate nonlinear functions. A single-layer decoupling can be viewed as a fully connecte…
Neural Network CompressionModel CompressionNeural Network Pruning via QUBO Optimization
Neural network pruning can be formulated as a combinatorial optimization problem, yet most existing approaches rely on greedy heuristics that ignore complex interactions between filters. Formal optimization methods such …
Neural Network CompressionImage DenoisingNetwork PruningPrune-Quantize-Distill: An Ordered Pipeline for Efficient Neural Network Compression
Modern deployment often requires trading accuracy for efficiency under tight CPU and memory constraints, yet common compression proxies such as parameter count or FLOPs do not reliably predict wall-clock inference time. …
Neural Network CompressionKnowledge DistillationAsymptotic Optimism for Tensor Regression Models with Applications to Neural Network Compression
We study rank selection for low-rank tensor regression under random covariates design. Under a Gaussian random-design model and some mild conditions, we derive population expressions for the expected training-testing dis…
Neural Network CompressionSimCert: Probabilistic Certification for Behavioral Similarity in Deep Neural Network Compression
Deploying Deep Neural Networks (DNNs) on resource-constrained embedded systems requires aggressive model compression techniques like quantization and pruning. However, ensuring that the compressed model preserves the beh…
Neural Network CompressionModel CompressionA Benchmark Study of Neural Network Compression Methods for Hyperspectral Image Classification
Deep neural networks have achieved strong performance in image classification tasks due to their ability to learn complex patterns from high-dimensional data. However, their large computational and memory requirements of…
Hyperspectral Image ClassificationNeural Network CompressionKnowledge DistillationAgenticPruner: MAC-Constrained Neural Network Compression via LLM-Driven Strategy Search
Neural network pruning remains essential for deploying deep learning models on resource-constrained devices, yet existing approaches primarily target parameter reduction without directly controlling computational cost. T…
Neural Network CompressionNetwork PruningConcatenated Matrix SVD: Compression Bounds, Incremental Approximation, and Error-Constrained Clustering
Large collections of matrices arise throughout modern machine learning, signal processing, and scientific computing, where they are commonly compressed by concatenation followed by truncated singular value decomposition …
Neural Network CompressionLow-Rank Prehab: Preparing Neural Networks for SVD Compression
Low-rank approximation methods such as singular value decomposition (SVD) and its variants (e.g., Fisher-weighted SVD, Activation SVD) have recently emerged as effective tools for neural network compression. In this sett…
Neural Network CompressionIDAP++: Advancing Divergence-Based Pruning via Filter-Level and Layer-Level Optimization
This paper presents a novel approach to neural network compression that addresses redundancy at both the filter and architectural levels through a unified framework grounded in information flow analysis. Building on the …
Neural Network CompressionModel CompressionPruning and Quantization Impact on Graph Neural Networks
Graph neural networks (GNNs) are known to operate with high accuracy on learning from graph-structured data, but they suffer from high computational and resource costs. Neural network compression methods are used to redu…
Neural Network CompressionGraph ClassificationNode ClassificationLink PredictionC-SWAP: Explainability-Aware Structured Pruning for Efficient Neural Networks Compression
Neural network compression has gained increasing attention in recent years, particularly in computer vision applications, where the need for model reduction is crucial for overcoming deployment constraints. Pruning is a …
Neural Network CompressionBinary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix Compression
This paper proposes a novel matrix quantization method, Binary Quadratic Quantization (BQQ). In contrast to conventional first-order quantization approaches, such as uniform quantization and binary coding quantization, t…
Neural Network Compression