Blackbird's language matrices (BLMs): a new benchmark to investigate disentangled generalisation in neural networks
Current successes of machine learning architectures are based on computationally expensive algorithms and prohibitively large amounts of data. We need to develop tasks and data to train networks to reach more complex and more compositional skills. In this paper, we illustrate Blackbird's language matrices (BLMs), a novel grammatical dataset developed to test a linguistic variant of Raven's progressive matrices, an intelligence test usually based on visual stimuli. The dataset consists of roughly 48000 sentences, generatively constructed to support investigations of current models' linguistic mastery of grammatical rules and their ability to generalize them. We present the logic of the dataset, the method to automatically construct data on a large scale, and the architecture to learn them. Through error analysis and several experiments on variations of the dataset, we demonstrate that this language task and the data that instantiate it provide a new challenging testbed to understand generalization and abstraction.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Blackbird's language matrices (BLMs): a new benchmark to investigate disentangled generalisation in neural networks
Current successes of machine learning architectures are based on computationally expensive algorithms and prohibitively large amounts of data. We need to develop tasks and data to train networks to reach more complex and…
Blackbird Language Matrices: A Framework to Investigate the Linguistic Competence of Language Models
This article describes a novel language task, the Blackbird Language Matrices (BLM) task, inspired by intelligence tests, and illustrates the BLM datasets, their construction and benchmarking, and targeted experiments on…
Exploring Italian sentence embeddings properties through multi-tasking
We investigate to what degree existing LLMs encode abstract linguistic information in Italian in a multi-task setting. We exploit curated synthetic data on a large scale -- several Blackbird Language Matrices (BLMs) prob…
SentenceSentence EmbeddingsExploring syntactic information in sentence embeddings through multilingual subject-verb agreement
In this paper, our goal is to investigate to what degree multilingual pretrained language models capture cross-linguistically valid abstract linguistic representations. We take the approach of developing curated syntheti…
Multiple-choiceSentenceSentence EmbeddingsvalidComparing Natural and Synthetic Structured Data: A Study of the Passive Verb Alternation in French and Italian
This study compares the impact of natural and synthetic data on training and evaluating large language models (LLMs), using the case of passive verb alternation in French and Italian. We use Blackbird Language Matrices (…