paper-with-me

홈 › Papers

State-Reification Networks: Improving Generalization by Modeling the Distribution of Hidden Representations

2019-05-26 · Alex Lamb, Jonathan Binas, Anirudh Goyal, Sandeep Subramanian, Ioannis Mitliagkas, Denis Kazakov, Yoshua Bengio, Michael C. Mozer

Machine learning promises methods that generalize well from finite labeled data. However, the brittleness of existing neural net approaches is revealed by notable failures, such as the existence of adversarial examples that are misclassified despite being nearly identical to a training example, or the inability of recurrent sequence-processing nets to stay on track without teacher forcing. We introduce a method, which we refer to as \emph{state reification}, that involves modeling the distribution of hidden states over the training data and then projecting hidden states observed during testing toward this distribution. Our intuition is that if the network can remain in a familiar manifold of hidden space, subsequent layers of the net should be well trained to respond appropriately. We show that this state-reification method helps neural nets to generalize better, especially when labeled data are sparse, and also helps overcome the challenge of achieving robust generalization with adversarial training.

📄 PDF Abstract BibTeX arXiv:1905.11382

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Probabilistic modeling the hidden layers of deep neural networks

2019-09-25 · Xinjie Lan, Kenneth E. Barner

In this paper, we demonstrate that the parameters of Deep Neural Networks (DNNs) cannot satisfy the i.i.d. prior assumption and activations being i.i.d. is not valid for all the hidden layers of DNNs. Hence, the Gaussian…

valid

Using Multiple Samples to Learn Mixture Models

2013-11-28 · NeurIPS 2013 12 · Jason D. Lee, Ran Gilad-Bachrach, Rich Caruana

In the mixture models problem it is assumed that there are $K$ distributions $\theta_{1},\ldots,\theta_{K}$ and one gets to observe a sample from a mixture of these distributions with unknown coefficients. The goal is to…

Learning Tree Distributions by Hidden Markov Models

2018-05-31 · Davide Bacciu, Daniele Castellana

Hidden tree Markov models allow learning distributions for tree structured data while being interpretable as nondeterministic automata. We provide a concise summary of the main approaches in literature, focusing in parti…

HiddenCut: Simple Data Augmentation for Natural Language Understanding with Better Generalization

2021-05-31 · Jiaao Chen, Dinghan Shen, Weizhu Chen, Diyi Yang

Fine-tuning large pre-trained models with task-specific data has achieved great success in NLP. However, it has been demonstrated that the majority of information within the self-attention networks is redundant and not u…

Data AugmentationNatural Language Understanding

Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

2024-06-14 · Rui Yang, Ruomeng Ding, Yong Lin, huan zhang 외

Reward models trained on human preference data have been proven to effectively align Large Language Models (LLMs) with human intent within the framework of reinforcement learning from human feedback (RLHF). However, curr…

Language ModelingLanguage ModellingText Generation