paper-with-me

홈 › Papers

PUMA: margin-based data pruning

2024-05-10 · Javier Maroto, Pascal Frossard

Deep learning has been able to outperform humans in terms of classification accuracy in many tasks. However, to achieve robustness to adversarial perturbations, the best methodologies require to perform adversarial training on a much larger training set that has been typically augmented using generative models (e.g., diffusion models). Our main objective in this work, is to reduce these data requirements while achieving the same or better accuracy-robustness trade-offs. We focus on data pruning, where some training samples are removed based on the distance to the model classification boundary (i.e., margin). We find that the existing approaches that prune samples with low margin fails to increase robustness when we add a lot of synthetic data, and explain this situation with a perceptron learning task. Moreover, we find that pruning high margin samples for better accuracy increases the harmful impact of mislabeled perturbed data in adversarial training, hurting both robustness and accuracy. We thus propose PUMA, a new data pruning strategy that computes the margin using DeepFool, and prunes the training samples of highest margin without hurting performance by jointly adjusting the training attack norm on the samples of lowest margin. We show that PUMA can be used on top of the current state-of-the-art methodology in robustness, and it is able to significantly improve the model performance unlike the existing data pruning strategies. Not only PUMA achieves similar robustness with less data, but it also significantly increases the model accuracy, improving the performance trade-off.

📄 PDF Abstract BibTeX arXiv:2405.06298

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Pruning 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Focus 설명 없음

Similar Papers 제목 키워드 기반

Update 3.0 to “PuMA: The Porous Microstructure Analysis software”

2021-07-19 · SoftwareX 2021 7 · Joseph C. Ferguson, Federico Semeraro, John M. Thornton, Francesco Panerai 외

A major update of the Porous Microstructure Analysis (PuMA) software is presented. PuMA is a framework for computing effective material properties and response based on material microstructures. Version 3.0 of the softwa…

Computed Tomography (CT)Physical Simulations

PUMA: Empowering Unified MLLM with Multi-granular Visual Generation

2024-10-17 · Rongyao Fang, Chengqi Duan, Kun Wang, Hao Li 외

Recent advancements in multimodal foundation models have yielded significant progress in vision-language understanding. Initial attempts have also explored the potential of multimodal large language models (MLLMs) for vi…

DiversityImage GenerationImage ManipulationText to Image Generation+1

PUMA: Performance Unchanged Model Augmentation for Training Data Removal

2022-03-02 · Ga Wu, Masoud Hashemi, Christopher Srinivasa

Preserving the performance of a trained model while removing unique characteristics of marked training data points is challenging. Recent research usually suggests retraining a model from scratch with remaining training …

Model Optimization

A Practitioner's Guide to Bayesian Inference in Pharmacometrics using Pumas

2023-03-31 · Mohamed Tarek, Jose Storopoli, Casey Davis, Chris Elrod 외

This paper provides a comprehensive tutorial for Bayesian practitioners in pharmacometrics using Pumas workflows. We start by giving a brief motivation of Bayesian inference for pharmacometrics highlighting limitations i…

Bayesian Inference

Pathway-Activity Likelihood Analysis and Metabolite Annotation for Untargeted Metabolomics using Probabilistic Modeling

2019-12-12 · Ramtin Hosseini, Neda Hassanpour, Li-Ping Liu, Soha Hassoun

Motivation: Untargeted metabolomics comprehensively characterizes small molecules and elucidates activities of biochemical pathways within a biological sample. Despite computational advances, interpreting collected measu…