paper-with-me

Papers

CheeseBench: Evaluating Large Language Models on Rodent Behavioral Neuroscience Paradigms

2026-04-12 · Zacharie Bugaud arxiv

We introduce CheeseBench, a benchmark that evaluates large language models (LLMs) on nine classical behavioral neuroscience paradigms (Morris water maze, Barnes maze, T-maze, radial arm maze, star maze, operant chamber, shuttle box, conditioned place preference, and delayed non-match to sample), spanning six cognitive dimensions. Each task is grounded in peer-reviewed rodent protocols with approximate animal baselines. The agent receives a unified system prompt with no task-specific instructions and must discover goals purely from ASCII text observations and reward signals, much like a rodent placed into an unfamiliar apparatus. We evaluate six open-weight LLMs (3B to 72B parameters) on text-based ASCII renderings and compare against both a random baseline and a graph-based reinforcement learning agent. Our best model (Qwen2.5-VL-7B) reaches 52.6% average success on ASCII input, compared to 32.1% for random agents and 78.9% for approximate rodent baselines. We find that (1) scaling beyond 7B yields diminishing returns, (2) longer context history degrades performance, (3) chain-of-thought prompting hurts rather than helps, and (4) a vision-language architecture provides an advantage at 7B but hurts at 32B. Because the same model's performance ranges from 20% to 57% depending on interface parameters alone, these results characterize the agent-plus-interface system, not the model in isolation. Under this unified zero-shot ASCII protocol, current open-weight LLM agents remain well below approximate rodent reference values, particularly on tasks requiring spatial navigation and within-trial state tracking.

📄 PDF Abstract BibTeX arXiv:2604.10825

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Rodent-Bench

2026-02-20 · Thomas Heap, Laurence Aitchison, Emma Cahill, Adriana Casado Rodriguez arxiv

We present Rodent-Bench, a novel benchmark designed to evaluate the ability of Multimodal Large Language Models (MLLMs) to annotate rodent behaviour footage. We evaluate state-of-the-art MLLMs, including Gemini-2.5-Pro, …

Data Augmentation for Automated Adaptive Rodent Training

2024-10-23 · Dibyendu Das, Alfredo Fontanini, Joshua F. Kogan, Haibin Ling 외

Fully optimized automation of behavioral training protocols for lab animals like rodents has long been a coveted goal for researchers. It is an otherwise labor-intensive and time-consuming process that demands close inte…

Data Augmentation

High-Precision UWB-Based Real-Time Locating System for Rodent Behavioral Studies

2024-09-03 · Reza Sayfoori, Mao-Hsiang Huang, Amir Naderi, Mehwish Bhatti 외

Rodents have long been established as the premier model for behavioral studies, traditionally raised and maintained in conventional cage environments. However, these settings often limit rodents' ability to exhibit their…

Deep neuroethology of a virtual rodent

2019-11-21 · ICLR 2020 1 · Josh Merel, Diego Aldarondo, Jesse Marshall, Yuval Tassa 외

Parallel developments in neuroscience and deep learning have led to mutually productive exchanges, pushing our understanding of real and artificial neural networks in sensory and cognitive systems. However, this interact…

Deep Reinforcement Learning

Evaluating U-net Brain Extraction for Multi-site and Longitudinal Preclinical Stroke Imaging

2022-03-11 · Erendiz Tarakci, Joseph Mandeville, Fahmeed Hyder, Basavaraju G. Sanganahalli 외

Rodent stroke models are important for evaluating treatments and understanding the pathophysiology and behavioral changes of brain ischemia, and magnetic resonance imaging (MRI) is a valuable tool for measuring outcome i…