fMRI predictors based on language models of increasing complexity recover brain left lateralization
Over the past decade, studies of naturalistic language processing where participants are scanned while listening to continuous text have flourished. Using word embeddings at first, then large language models, researchers have created encoding models to analyze the brain signals. Presenting these models with the same text as the participants allows to identify brain areas where there is a significant correlation between the functional magnetic resonance imaging (fMRI) time series and the ones predicted by the models' artificial neurons. One intriguing finding from these studies is that they have revealed highly symmetric bilateral activation patterns, somewhat at odds with the well-known left lateralization of language processing. Here, we report analyses of an fMRI dataset where we manipulate the complexity of large language models, testing 28 pretrained models from 8 different families, ranging from 124M to 14.2B parameters. First, we observe that the performance of models in predicting brain responses follows a scaling law, where the fit with brain activity increases linearly with the logarithm of the number of parameters of the model (and its performance on natural language processing tasks). Second, although this effect is present in both hemispheres, it is stronger in the left than in the right hemisphere. Specifically, the left-right difference in brain correlation follows a scaling law with the number of parameters. This finding reconciles computational analyses of brain activity using large language models with the classic observation from aphasic patients showing left hemisphere dominance for language.
Code (1)
Tasks
Word EmbeddingsSimilar Papers 제목 키워드 기반
The Alice Datasets: fMRI \& EEG Observations of Natural Language Comprehension
The Alice Datasets are a set of datasets based on magnetic resonance data and electrophysiological data, collected while participants heard a story in English. Along with the datasets and the text of the story, we provid…
EEGElectroencephalogram (EEG)validMore Than Meets the Eye: Self-Supervised Depth Reconstruction From Brain Activity
In the past few years, significant advancements were made in reconstruction of observed natural images from fMRI brain recordings using deep-learning tools. Here, for the first time, we show that dense 3D depth maps of o…
Brain DecodingDiscovering Dynamic Functional Brain Networks via Spatial and Channel-wise Attention
Using deep learning models to recognize functional brain networks (FBNs) in functional magnetic resonance imaging (fMRI) has been attracting increasing interest recently. However, most existing work focuses on detecting …
Functional ConnectivityMapping fNIRS to fMRI with Neural Data Augmentation and Machine Learning Models
Advances in neuroimaging techniques have provided us novel insights into understanding how the human mind works. Functional magnetic resonance imaging (fMRI) is the most popular and widely used neuroimaging technique, an…
BIG-bench Machine LearningData AugmentationSNAKE-fMRI: A modular fMRI data simulator from the space-time domain to k-space and back
We propose a new, modular, open-source, Python-based 3D+time fMRI data simulation software, \emph{SNAKE-fMRI}, which stands for \emph{S}imulator from \emph{N}eurovascular coupling to \emph{A}cquisition of \emph{K}-space …
Image Reconstruction