paper-with-me

Papers Neural Network Compression

“Neural Network Compression” 태그가 달린 논문 215편 · 필터 해제

M-Fibration Theory with Applications to Weighted Graphs

2026-08-26 · Paolo Boldi arxiv

The purpose of this paper is to provide a general, comprehensive, theoretical framework that allows one to deal with fibrations on graphs labelled on a commutative monoid. This is a genuine extension of the theory of gra…

Neural Network Compression

On the Applicability of Safety Nets: A Safety-By-Design Solution for Certifying Neural Networks

2026-08-20 · Johann Maximilian Christensen, Thomas Stefani, Elena Hoemann, Frank Köster 외 arxiv

The integration of Artificial Intelligence (AI) in safety-critical aviation systems presents significant challenges for certification and deployment. Aviation, often regarded as the safest form of transportation, relies …

Neural Network Compression

EvoLP: Self-Evolving Latency Predictor for Model Compression in Real-Time Edge Systems

2026-07-10 · Shuo Huai, Hao Kong, Shiqing Li, Xiangzhong Luo 외 arxiv

Edge devices are increasingly utilized for deploying deep learning applications on embedded systems. The real-time nature of many applications and the limited resources of edge devices necessitate latency-targeted neural…

Neural Network CompressionModel Compression

Hierarchical Reinforcement Learning for Neural Network Compression (HiReLC): Pruning and Quantization

2026-06-24 · Kamar Hibatallah Baghdadi, Kawther Guoual Belhamidi, Sara Belhadj, Aissa Boulmerka 외 arxiv

We present HiReLC, a hierarchical ensemble-reinforcement learning framework for automated joint quantization and structured pruning of deep neural networks. The framework decomposes the compression search across two leve…

Hierarchical Reinforcement LearningNeural Network CompressionActive Learning

Hybrid Compression: Integrating Pruning and Quantization for Optimized Neural Networks

2026-06-22 · Minh-Loi Nguyen, Long-Bao Nguyen, Van-Hieu Huynh, Minh-Triet Tran 외 arxiv

Deep neural networks have witnessed remarkable advancements in recent years and have become integral to various applications. However, alongside these developments, training and deployment of neural network models on emb…

Neural Network CompressionModel Compression

Neural Network Compression by Approximate Differential Equivalence

2026-05-31 · Ravi Dhiman, Andrea Passarella, Mirco Tribastone, Lorenzo Valerio arxiv

Neural network compression is commonly achieved by pruning parameters based on local importance scores, e.g., magnitude-based pruning. We propose a complementary approach that compresses models by aggregating neurons wit…

Neural Network Compression

Fast Tensorization of Neural Networks via Slice-wise Feature Distillation

2026-05-19 · Safa Hamreras, Sukhbinder Singh, Román Orús arxiv

We propose a scalable tensorization framework for neural network compression based on slice-wise feature distillation. Unlike conventional tensor decomposition methods that rely on costly global finetuning, our approach …

Neural Network Compression

Robust Basis Spline Decoupling for the Compression of Transformer Models

2026-05-11 · Joppe De Jonghe, Van Tien Pham, Mariya Ishteva arxiv

Decoupling is a powerful modeling paradigm for representing multivariate functions as compositions of linear transformations and univariate nonlinear functions. A single-layer decoupling can be viewed as a fully connecte…

Neural Network CompressionModel Compression

Neural Network Pruning via QUBO Optimization

2026-04-07 · Osama Orabi, Artur Zagitov, Hadi Salloum, Viktor A. Lobachev 외 arxiv

Neural network pruning can be formulated as a combinatorial optimization problem, yet most existing approaches rely on greedy heuristics that ignore complex interactions between filters. Formal optimization methods such …

Neural Network CompressionImage DenoisingNetwork Pruning

Prune-Quantize-Distill: An Ordered Pipeline for Efficient Neural Network Compression

2026-04-05 · Longsheng Zhou, Yu Shen arxiv

Modern deployment often requires trading accuracy for efficiency under tight CPU and memory constraints, yet common compression proxies such as parameter count or FLOPs do not reliably predict wall-clock inference time. …

Neural Network CompressionKnowledge Distillation

Asymptotic Optimism for Tensor Regression Models with Applications to Neural Network Compression

2026-03-27 · Haoming Shi, Eric C. Chi, Hengrui Luo arxiv

We study rank selection for low-rank tensor regression under random covariates design. Under a Gaussian random-design model and some mild conditions, we derive population expressions for the expected training-testing dis…

Neural Network Compression

SimCert: Probabilistic Certification for Behavioral Similarity in Deep Neural Network Compression

2026-03-16 · Jingyang Li, Fu Song, Guoqiang Li arxiv

Deploying Deep Neural Networks (DNNs) on resource-constrained embedded systems requires aggressive model compression techniques like quantization and pruning. However, ensuring that the compressed model preserves the beh…

Neural Network CompressionModel Compression

A Benchmark Study of Neural Network Compression Methods for Hyperspectral Image Classification

2026-03-05 · Sai Shi arxiv

Deep neural networks have achieved strong performance in image classification tasks due to their ability to learn complex patterns from high-dimensional data. However, their large computational and memory requirements of…

Hyperspectral Image ClassificationNeural Network CompressionKnowledge Distillation

AgenticPruner: MAC-Constrained Neural Network Compression via LLM-Driven Strategy Search

2026-01-18 · Shahrzad Esmat, Mahdi Banisharif, Ali Jannesari arxiv

Neural network pruning remains essential for deploying deep learning models on resource-constrained devices, yet existing approaches primarily target parameter reduction without directly controlling computational cost. T…

Neural Network CompressionNetwork Pruning

Concatenated Matrix SVD: Compression Bounds, Incremental Approximation, and Error-Constrained Clustering

2026-01-12 · Maksym Shamrai arxiv

Large collections of matrices arise throughout modern machine learning, signal processing, and scientific computing, where they are commonly compressed by concatenation followed by truncated singular value decomposition …

Neural Network Compression

Low-Rank Prehab: Preparing Neural Networks for SVD Compression

2025-12-01 · Haoran Qin, Shansita Sharma, Ali Abbasi, Chayne Thrash 외 arxiv

Low-rank approximation methods such as singular value decomposition (SVD) and its variants (e.g., Fisher-weighted SVD, Activation SVD) have recently emerged as effective tools for neural network compression. In this sett…

Neural Network Compression

IDAP++: Advancing Divergence-Based Pruning via Filter-Level and Layer-Level Optimization

2025-11-25 · Aleksei Samarin, Artem Nazarenko, Egor Kotenko, Valentin Malykh 외 arxiv

This paper presents a novel approach to neural network compression that addresses redundancy at both the filter and architectural levels through a unified framework grounded in information flow analysis. Building on the …

Neural Network CompressionModel Compression

Pruning and Quantization Impact on Graph Neural Networks

2025-10-24 · Khatoon Khedri, Reza Rawassizadeh, Qifu Wen, Mehdi Hosseinzadeh arxiv

Graph neural networks (GNNs) are known to operate with high accuracy on learning from graph-structured data, but they suffer from high computational and resource costs. Neural network compression methods are used to redu…

Neural Network CompressionGraph ClassificationNode ClassificationLink Prediction

C-SWAP: Explainability-Aware Structured Pruning for Efficient Neural Networks Compression

2025-10-21 · Baptiste Bauvin, Loïc Baret, Ola Ahmad arxiv

Neural network compression has gained increasing attention in recent years, particularly in computer vision applications, where the need for model reduction is crucial for overcoming deployment constraints. Pruning is a …

Neural Network Compression

Binary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix Compression

2025-10-21 · Kyo Kuroki, Yasuyuki Okoshi, Thiem Van Chu, Kazushi Kawamura 외 arxiv

This paper proposes a novel matrix quantization method, Binary Quadratic Quantization (BQQ). In contrast to conventional first-order quantization approaches, such as uniform quantization and binary coding quantization, t…

Neural Network Compression
1–20 / 215 다음 →