paper-with-me

Model Compression

2개 벤치마크 · 논문 1,630편 · 이 태스크의 논문 보기 →

Benchmarks

ImageNet

결과 14개

QNLI

결과 2개

Most implemented

Learned Step Size Quantization

2019-02-21 · 구현 9개

Papers

Debias-SparseGPT: Bias-Aware Pruning for Large Language Models

2026-09-02 · Irina Proskurina, Guillaume Metzler, Antoine Gourru, Julien Velcin hf

Model compression techniques such as pruning and quantization facilitate the efficient deployment and acceleration of Large Language Models (LLMs). However, recent studies show that weight sparsification methods, such as…

Computational EfficiencyModel Compression

When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs

2026-08-26 · Suchit Gupte, Xueru Zhang, Mohammad Mahdi Khalili arxiv

Sparse autoencoders (SAEs) are widely used to interpret the internal representations of large language models (LLMs), yet their reliability under post-hoc model compression remains poorly understood. We present a systema…

Model Compression

Exploring the Performance Frontier of Compact Unified Image Generation Models

2026-08-20 · Taihang Hu, Zhao Wang, Zuan Gao, Tao Liu 외 arxiv

We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing. Our goal is to explore how far a relatively small visual generator can be pushed through system…

Text-to-Image GenerationReinforcement LearningModel CompressionImage Editing

Kilobyte Models: Neural Networks as a Seed and a Quantized Latent

2026-08-01 · Sahil Rajesh Dhayalkar arxiv

The cost of storing and transmitting a trained neural network scales with its parameter count, a bottleneck for over-the-air updates, on-device libraries, and other bandwidth-bound deployments. We study an extreme form o…

Model Compression

Memory Efficient Tabular Foundation Models

2026-07-30 · Shuting Luo, Monika Mikhail Kanaan, Cameron Gordon, Anna Leontjeva 외 arxiv

Tabular Foundation Models, such as TabPFN, have received a large amount of recent attention due to their performance on in-context tabular machine learning tasks, which often exceeds classical baselines. However, practic…

Model Compression

CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation

2026-07-29 · Fengming Yu, Haiwei Pan, Kejia Zhang, Chunling Chen 외 arxiv

Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. The emergence of diverse model architectures has extended KD from h…

Knowledge DistillationModel Compression

전체 1,630편 보기 →