paper-with-me

홈 › Papers

Frequency Matters: Fast Model-Agnostic Data Curation for Pruning and Quantization

2026-03-17 · Francesco Pio Monaco, Elia Cunegatti, Flavio Vella, Giovanni Iacca arxiv

Post-training model compression is essential for enhancing the portability of Large Language Models (LLMs) while preserving their performance. While several compression approaches have been proposed, less emphasis has been placed on selecting the most suitable set of data (the so-called \emph{calibration data}) for finding the compressed model configuration. The choice of calibration data is a critical step in preserving model capabilities both intra- and inter-tasks. In this work, we address the challenge of identifying high-performance calibration sets for both pruning and quantization by analyzing intrinsic data properties rather than model-specific signals. We introduce ZipCal, a model-agnostic data curation strategy that maximizes lexical diversity based on Zipfian power laws. Experiments demonstrate that our method outperforms standard uniform random sampling across various pruning benchmarks. Notably, it also performs on par, in terms of downstream performance, with a state-of-the-art method that relies on model perplexity. The latter becomes prohibitively expensive for large-scale models and datasets, while ZipCal is on average $\sim$240$\times$ faster due to its tractable linear complexity. We make the code and the experiments available at https://github.com/FrancescoMonaco/ZipCal.

📄 PDF Abstract BibTeX arXiv:2603.16105

Code (0)

등록된 구현이 없습니다.

Tasks

Model Compression

Similar Papers 제목 키워드 기반

A Survey on Data Curation for Visual Contrastive Learning: Why Crafting Effective Positive and Negative Pairs Matters

2025-02-12 · Shasvat Desai, Debasmita Ghose, Deep Chakraborty

Visual contrastive learning aims to learn representations by contrasting similar (positive) and dissimilar (negative) pairs of data samples. The design of these pairs significantly impacts representation quality, trainin…

Contrastive Learning

Concept-Aware Batch Sampling Improves Language-Image Pretraining

2025-11-25 · Adhiraj Ghosh, Vishaal Udandarao, Thao Nguyen, Matteo Farina 외 arxiv

What data should a vision-language model be trained on? To answer this question, many data curation efforts center on the quality of a dataset. However, most of these existing methods are (i) offline, i.e. they produce a…

fastMRI Breast: A publicly available radial k-space dataset of breast dynamic contrast-enhanced MRI

2024-06-07 · Eddy Solomon, Patricia M. Johnson, Zhengguo Tan, Radhika Tibrewala 외

This data curation work introduces the first large-scale dataset of radial k-space and DICOM data for breast DCE-MRI acquired in diagnostic breast MRI exams. Our dataset includes case-level labels indicating patient age,…

DiagnosticImage Reconstruction

What Matters in Data Curation for Multimodal Reasoning? Insights from the DCVLR Challenge

2026-01-16 · Yosub Shin, Michael Buriek, Boris Sobolev, Pavel Bushuyeu 외 arxiv

We study data curation for multimodal reasoning through the NeurIPS 2025 Data Curation for Vision-Language Reasoning (DCVLR) challenge, which isolates dataset selection by fixing the model and training protocol. Using a …

Multimodal Reasoning

Correcting Neural Operator Spectral Bias via Diffusion Posterior Sampling with Sparse Observations

2026-06-02 · Niccolò Perrone, Fanny Lehmann, Stefania Fresca, Filippo Gatti arxiv

Neural operator surrogates (NO) approximate PDE solutions orders of magnitude faster than numerical solvers, but suffer from spectral bias: high-frequency content is systematically attenuated, limiting reliability where …