paper-with-me

홈 › Papers

A global predicted-fMRI drive signal from TRIBE does not predict YouTube replay heatmaps

2026-07-01 · Barada Sahu, Shivesh Pandey arxiv

Deep multimodal brain-encoding models now predict fMRI responses to naturalistic video with high accuracy; whether their predicted neural signals also forecast behavioral engagement is unknown. We run TRIBE, the winning model of the 2025 Algonauts challenge (Llama-3.2 + V-JEPA 2 + Wav2Vec-BERT), on 48 YouTube videos and reduce its predicted cortical response to a per-second engagement curve, the global field power. Correlated against each video's "most replayed" heatmap, a proxy for re-watch, it shows no evidence of prediction: the pooled position-controlled partial correlation is +0.058 (95% CI [-0.04, 0.15]; t(47)=1.21, p=0.23), and not above simple loudness/motion baselines. The raw correlation is also near zero; the moderate values for music videos are an onset-replay artifact. The null holds across six cortical-network readouts, value/salience ROIs, and a permutation test; a supervised leave-one-video-out probe appears to reach r=0.47 but collapses to a temporal-shape artifact under a proper position control. Running the probe on TRIBE's input streams reveals at most a small, borderline visual-stream signal (matched vs. mismatched p=0.004-0.06) and none in audio, text, or the predicted cortex. The inter-subject-correlation readout, the closest prior positive result, is unavailable from the subject-averaged released model, so we fit our own per-subject encoders on the Algonauts fMRI (validated in-domain at r=0.15 and cross-domain, Friends-to-film, at r=0.10); the predicted ISC still does not track re-watch (r=-0.04, p=0.34). We bound rather than merely fail to reject the null: a Bayes factor gives moderate evidence for it (BF01=3.2), an equivalence test excludes effects above r=0.14, and the target's split-half reliability (0.82; ceiling r=0.9) rules out a noisy-label artifact. We release code, a video-ID manifest, and a heatmap-acquisition method robust to YouTube's SABR streaming.

📄 PDF Abstract BibTeX arXiv:2607.01400

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Feature Visualization Recovers Known Cortical Selectivity from TRIBE v2

2026-05-13 · Stuart Bladon, Brinnae Bent arxiv

Brain encoder models predict cortical fMRI responses from the internal activations of pretrained vision and language networks, and are typically evaluated by held-out prediction accuracy. This is a useful signal for trai…

Boosting Brain-to-Image Decoding with TRIBE v2 Data Augmentation

2026-06-04 · Yohann Benchetrit, Marlène Careil, Simon Dahan, Hubert Banville 외 arxiv

Brain decoding is limited by the availability of labeled neural data, and remains challenging in low-data regimes. To address this issue, we investigate whether and when brain decoding can be boosted by augmenting small …

Data AugmentationBrain Decoding

TRIBE: TRImodal Brain Encoder for whole-brain fMRI response prediction

2025-07-29 · Stéphane d'Ascoli, Jérémy Rapin, Yohann Benchetrit, Hubert Banville 외 arxiv

Historically, neuroscience has progressed by fragmenting into specialized domains, each focusing on isolated modalities, tasks, or brain regions. While fruitful, this approach hinders the development of a unified model o…

A foundation model of vision, audition, and language for in-silico neuroscience

2026-05-05 · Stéphane d'Ascoli, Jérémy Rapin, Yohann Benchetrit, Teon Brooks 외 arxiv

Cognitive neuroscience is fragmented into specialized models, each tailored to specific experimental paradigms, hence preventing a unified model of cognition in the human brain. Here, we introduce TRIBE v2, a tri-modal (…

Predicted Cortex Is Not a Domain-General Prior: A Matched-Control Audit of Brain-Encoding Features for Video Memorability

2026-07-12 · Carson Rodrigues arxiv

Brain-encoding foundation models predict fMRI responses to video, audio and text well enough to win the Algonauts 2025 challenge. We ask whether their predicted responses, obtained with no scanner, are a useful feature l…