paper-with-me

홈 › Papers

Weight-Covariance Alignment for Adversarially Robust Neural Networks

2020-10-17 · Panagiotis Eustratiadis, Henry Gouk, Da Li, Timothy Hospedales

Stochastic Neural Networks (SNNs) that inject noise into their hidden layers have recently been shown to achieve strong robustness against adversarial attacks. However, existing SNNs are usually heuristically motivated, and often rely on adversarial training, which is computationally costly. We propose a new SNN that achieves state-of-the-art performance without relying on adversarial training, and enjoys solid theoretical justification. Specifically, while existing SNNs inject learned or hand-tuned isotropic noise, our SNN learns an anisotropic noise distribution to optimize a learning-theoretic bound on adversarial robustness. We evaluate our method on a number of popular benchmarks, show that it can be applied to different architectures, and that it provides robustness to a variety of white-box and black-box attacks, while being simple and fast to train compared to existing alternatives.

📄 PDF Abstract BibTeX arXiv:2010.08852

Code (1)

peustr/WCA-net 공식 구현 pytorch

Tasks

Adversarial Robustness

Similar Papers 제목 키워드 기반

Uncovering Cross-Objective Interference in Multi-Objective Alignment

2026-02-06 · Yining Lu, Meng Jiang arxiv

We study a persistent failure mode in multi-objective alignment for large language models (LLMs): training improves performance on only a subset of objectives while causing others to degrade. We formalize this phenomenon…

A Commutator Framework for Selective Spectral Alignment in Deep Neural Networks

2026-08-24 · Kaj Nyström arxiv

We develop a finite-width geometric framework describing how learned feature geometries are organized, transported, and selectively aligned in deep neural networks. Incompatibility among weight-generated covariance, gate…

Agnostic Estimation of Mean and Covariance

2016-04-24 · Kevin A. Lai, Anup B. Rao, Santosh Vempala

We consider the problem of estimating the mean and covariance of a distribution from iid samples in $\mathbb{R}^n$, in the presence of an $\eta$ fraction of malicious noise; this is in contrast to much recent work where …

AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss

2026-08-11 · Mingju Gao, Jingkai Zhou, Kun Gai, Changqian Yu 외 hf

Fréchet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-level diffusion and flow-matching losses. However, directly optimizing…

Where Pretraining writes and Alignment reads: the asymmetry of Transformer weight space

2026-05-15 · Valeria Ruscio, Eli-Shaoul Khedouri, Keiran Thompson arxiv

Cross-entropy pretraining and preference alignment update the same transformer weights, but leave geometrically distinct traces. We characterise this asymmetry with a relative-subspace-fraction probe that tracks how weig…