Generalized but not Robust? Comparing the Effects of Data Modification Methods on Out-of-Domain Generalization and Adversarial Robustness
Data modification, either via additional training datasets, data augmentation, debiasing, and dataset filtering, has been proposed as an effective solution for generalizing to out-of-domain (OOD) inputs, in both natural language processing and computer vision literature. However, the effect of data modification on adversarial robustness remains unclear. In this work, we conduct a comprehensive study of common data modification strategies and evaluate not only their in-domain and OOD performance, but also their adversarial robustness (AR). We also present results on a two-dimensional synthetic dataset to visualize the effect of each method on the training distribution. This work serves as an empirical study towards understanding the relationship between generalizing to unseen domains and defending against adversarial perturbations. Our findings suggest that more data (either via additional datasets or data augmentation) benefits both OOD accuracy and AR. However, data filtering (previously shown to improve OOD accuracy on natural language inference) hurts OOD accuracy on other tasks such as question answering and image classification. We provide insights from our experiments to inform future work in this direction.
Code (0)
등록된 구현이 없습니다.
Tasks
Adversarial RobustnessData AugmentationDomain Generalizationimage-classificationImage ClassificationNatural Language InferenceQuestion AnsweringSimilar Papers 제목 키워드 기반
Synthetic Controls with spillover effects: A comparative study
Iterative Synthetic Control Method is introduced in this study, a modification of the Synthetic Control Method (SCM) designed to improve its predictive performance by utilizing control units affected by the treatment in …
Simultaneous inference for generalized linear models with unmeasured confounders
Tens of thousands of simultaneous hypothesis tests are routinely performed in genomic studies to identify differentially expressed genes. However, due to unmeasured confounders, many standard statistical approaches may b…
Optimal Estimation of Generalized Average Treatment Effects using Kernel Optimal Matching
In causal inference, a variety of causal effect estimands have been studied, including the sample, uncensored, target, conditional, optimal subpopulation, and optimal weighted average treatment effects. Ad-hoc methods ha…
Causal InferenceGCF: Generalized Causal Forest for Heterogeneous Treatment Effect Estimation in Online Marketplace
Uplift modeling is a rapidly growing approach that utilizes causal inference and machine learning methods to directly estimate the heterogeneous treatment effects, which has been widely applied to various online marketpl…
Causal InferenceDecision MakingHeterogeneous Treatment Effect EstimationNeural stochastic Volterra equations: learning path-dependent dynamics
Stochastic Volterra equations (SVEs) serve as mathematical models for the time evolutions of random systems with memory effects and irregular behaviour. We introduce neural stochastic Volterra equations as a physics-insp…