paper-with-me

홈 › Papers

Invariant Language Modeling

2021-10-16 · Maxime Peyrard, Sarvjeet Singh Ghotra, Martin Josifoski, Vidhan Agarwal, Barun Patra, Dean Carignan, Emre Kiciman, Robert West

Large pretrained language models are critical components of modern NLP pipelines. Yet, they suffer from spurious correlations, poor out-of-domain generalization, and biases. Inspired by recent progress in causal machine learning, in particular the invariant risk minimization (IRM) paradigm, we propose invariant language modeling, a framework for learning invariant representations that generalize better across multiple environments. In particular, we adapt a game-theoretic formulation of IRM (IRM-games) to language models, where the invariance emerges from a specific training schedule in which all the environments compete to optimize their own environment-specific loss by updating subsets of the model in a round-robin fashion. We focus on controlled experiments to precisely demonstrate the ability of our method to (i) remove structured noise, (ii) ignore specific spurious correlations without affecting global performance, and (iii) achieve better out-of-domain generalization. These benefits come with a negligible computational overhead compared to standard training, do not require changing the local loss, and can be applied to any language model. We believe this framework is promising to help mitigate spurious correlations and biases in language models.

📄 PDF Abstract BibTeX arXiv:2110.08413

Code (1)

epfl-dlab/invariant-language-models 공식 구현 pytorch

Tasks

Domain GeneralizationLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Invariant Language Modeling

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Modern pretrained language models are critical components of NLP pipelines. Yet, they suffer from spurious correlations, poor out-of-domain generalization, and biases. Inspired by recent progress in causal machine learn…

Domain GeneralizationLanguage ModelingLanguage Modelling

Learning Language-Driven Sequence-Level Modal-Invariant Representations for Video-Based Visible-Infrared Person Re-Identification

2026-01-17 · Xiaomei Yang, Antai Liu, Xizhan Gao, Fa Zhu 외 arxiv

The core of video-based visible-infrared person re-identification (VVI-ReID) lies in learning sequence-level modal-invariant representations across different modalities. Recent research tends to use modality-shared langu…

Person Re-IdentificationRepresentation Learning

Adversarial Training for Multilingual Acoustic Modeling

2019-06-17 · Ke Hu, Hasim Sak, Hank Liao

Multilingual training has been shown to improve acoustic modeling performance by sharing and transferring knowledge in modeling different languages. Knowledge sharing is usually achieved by using common lower-level layer…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Identificationspeech-recognition+1

Importance of equivariant and invariant symmetries for fluid flow modeling

2023-05-03 · Varun Shankar, Shivam Barwey, Zico Kolter, Romit Maulik 외

Graph neural networks (GNNs) have shown promise in learning unstructured mesh-based simulations of physical systems, including fluid dynamics. In tandem, geometric deep learning principles have informed the development o…

ÚFAL at MRP 2020: Permutation-invariant Semantic Parsing in PERIN

2020-11-02 · David Samuel, Milan Straka

We present PERIN, a novel permutation-invariant approach to sentence-to-graph semantic parsing. PERIN is a versatile, cross-framework and language independent architecture for universal modeling of semantic structures. O…

Semantic ParsingSentence