paper-with-me

Papers

Systematicity Emerges in Transformers when Abstract Grammatical Roles Guide Attention

2022-07-01 · NAACL (ACL) 2022 7 · Ayush K Chakravarthy, Jacob Labe Russin, Randall O’Reilly

Systematicity is thought to be a key inductive bias possessed by humans that is lacking in standard natural language processing systems such as those utilizing transformers. In this work, we investigate the extent to which the failure of transformers on systematic generalization tests can be attributed to a lack of linguistic abstraction in its attention mechanism. We develop a novel modification to the transformer by implementing two separate input streams: a role stream controls the attention distributions (i.e., queries and keys) at each layer, and a filler stream determines the values. Our results show that when abstract role labels are assigned to input sequences and provided to the role stream, systematic generalization is improved.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Inductive BiasSystematic Generalization

Similar Papers 제목 키워드 기반

Systematicity between Forms and Meanings across Languages Supports Efficient Communication

2026-01-23 · Doreen Osmelak, Yang Xu, Michael Hahn, Kate McCurdy arxiv

Languages vary widely in how meanings map to word forms. These mappings have been found to support efficient communication; however, this theory does not account for systematic relations within word forms. We examine how…

Refining Targeted Syntactic Evaluation of Language Models

2021-04-19 · NAACL 2021 4 · Benjamin Newman, Kai-Siang Ang, Julia Gong, John Hewitt

Targeted syntactic evaluation of subject-verb number agreement in English (TSE) evaluates language models' syntactic knowledge using hand-crafted minimal pairs of sentences that differ only in the main verb's conjugation…

Sentence

Inducing Systematicity in Transformers by Attending to Structurally Quantized Embeddings

2024-02-09 · Yichen Jiang, Xiang Zhou, Mohit Bansal

Transformers generalize to novel compositions of structures and entities after being trained on a complex dataset, but easily overfit on datasets of insufficient complexity. We observe that when the training set is suffi…

Machine TranslationQuantizationSemantic ParsingWord Embeddings

Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language Models

2023-11-15 · James A. Michaelov, Catherine Arnett, Tyler A. Chang, Benjamin K. Bergen

Abstract grammatical knowledge - of parts of speech and grammatical patterns - is key to the capacity for linguistic generalization in humans. But how abstract is grammatical knowledge in large language models? In the hu…

Sentence

Explainability Through Systematicity: The Hard Systematicity Challenge for Artificial Intelligence

2025-07-29 · Matthieu Queloz arxiv

This paper argues that explainability is only one facet of a broader ideal that shapes our expectations towards artificial intelligence (AI). Fundamentally, the issue is to what extent AI exhibits systematicity--not mere…