paper-with-me

홈 › Papers

Blackbird's language matrices (BLMs): a new benchmark to investigate disentangled generalisation in neural networks

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Current successes of machine learning architectures are based on computationally expensive algorithms and prohibitively large amounts of data. We need to develop tasks and data to train networks to reach more complex and more compositional skills. In this paper, we illustrate Blackbird's language matrices (BLMs), a novel grammatical dataset developed to test a linguistic variant of Raven's progressive matrices, an intelligence test usually based on visual stimuli. The dataset consists of roughly 48000 sentences, generatively constructed to support investigations of current models' linguistic mastery of grammatical rules and their ability to generalize them. We present the logic of the dataset, the method to automatically construct data on a large scale, and the architecture to learn them. Through error analysis and several experiments on variations of the dataset, we demonstrate that this language task and the data that instantiate it provide a new challenging testbed to understand generalization and abstraction.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Blackbird's language matrices (BLMs): a new benchmark to investigate disentangled generalisation in neural networks

2022-05-22 · Paola Merlo, Aixiu An, Maria A. Rodriguez

Current successes of machine learning architectures are based on computationally expensive algorithms and prohibitively large amounts of data. We need to develop tasks and data to train networks to reach more complex and…

Blackbird Language Matrices: A Framework to Investigate the Linguistic Competence of Language Models

2026-02-24 · Paola Merlo, Chunyang Jiang, Giuseppe Samo, Vivi Nastase arxiv

This article describes a novel language task, the Blackbird Language Matrices (BLM) task, inspired by intelligence tests, and illustrates the BLM datasets, their construction and benchmarking, and targeted experiments on…

Exploring Italian sentence embeddings properties through multi-tasking

2024-09-10 · Vivi Nastase, Giuseppe Samo, Chunyang Jiang, Paola Merlo

We investigate to what degree existing LLMs encode abstract linguistic information in Italian in a multi-task setting. We exploit curated synthetic data on a large scale -- several Blackbird Language Matrices (BLMs) prob…

SentenceSentence Embeddings

Exploring syntactic information in sentence embeddings through multilingual subject-verb agreement

2024-09-10 · Vivi Nastase, Chunyang Jiang, Giuseppe Samo, Paola Merlo

In this paper, our goal is to investigate to what degree multilingual pretrained language models capture cross-linguistically valid abstract linguistic representations. We take the approach of developing curated syntheti…

Multiple-choiceSentenceSentence Embeddingsvalid

Comparing Natural and Synthetic Structured Data: A Study of the Passive Verb Alternation in French and Italian

2026-03-26 · Giuseppe Samo, Paola Merlo arxiv

This study compares the impact of natural and synthetic data on training and evaluating large language models (LLMs), using the case of passive verb alternation in French and Italian. We use Blackbird Language Matrices (…