paper-with-me

Papers

Rodent-Bench

2026-02-20 · Thomas Heap, Laurence Aitchison, Emma Cahill, Adriana Casado Rodriguez arxiv

We present Rodent-Bench, a novel benchmark designed to evaluate the ability of Multimodal Large Language Models (MLLMs) to annotate rodent behaviour footage. We evaluate state-of-the-art MLLMs, including Gemini-2.5-Pro, Gemini-2.5-Flash and Qwen-VL-Max, using this benchmark and find that none of these models perform strongly enough to be used as an assistant for this task. Our benchmark encompasses diverse datasets spanning multiple behavioral paradigms including social interactions, grooming, scratching, and freezing behaviors, with videos ranging from 10 minutes to 35 minutes in length. We provide two benchmark versions to accommodate varying model capabilities and establish standardized evaluation metrics including second-wise accuracy, macro F1, mean average precision, mutual information, and Matthew's correlation coefficient. While some models show modest performance on certain datasets (notably grooming detection), overall results reveal significant challenges in temporal segmentation, handling extended video sequences, and distinguishing subtle behavioral states. Our analysis identifies key limitations in current MLLMs for scientific video annotation and provides insights for future model development. Rodent-Bench serves as a foundation for tracking progress toward reliable automated behavioral annotation in neuroscience research.

📄 PDF Abstract BibTeX arXiv:2602.18540

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CheeseBench: Evaluating Large Language Models on Rodent Behavioral Neuroscience Paradigms

2026-04-12 · Zacharie Bugaud arxiv

We introduce CheeseBench, a benchmark that evaluates large language models (LLMs) on nine classical behavioral neuroscience paradigms (Morris water maze, Barnes maze, T-maze, radial arm maze, star maze, operant chamber, …

Reinforcement Learning

Data Augmentation for Automated Adaptive Rodent Training

2024-10-23 · Dibyendu Das, Alfredo Fontanini, Joshua F. Kogan, Haibin Ling 외

Fully optimized automation of behavioral training protocols for lab animals like rodents has long been a coveted goal for researchers. It is an otherwise labor-intensive and time-consuming process that demands close inte…

Data Augmentation

RodEpil: A Video Dataset of Laboratory Rodents for Seizure Detection and Benchmark Evaluation

2025-11-13 · Daniele Perlo, Vladimir Despotovic, Selma Boudissa, Sang-Yoon Kim 외 arxiv

We introduce a curated video dataset of laboratory rodents for automatic detection of convulsive events. The dataset contains short (10~s) top-down and side-view video clips of individual rodents, labeled at clip level a…

Seizure Detection

Formal Concept Analysis of Rodent Carriers of Zoonotic Disease

2016-08-25 · Roman Ilin, Barbara A. Han

The technique of Formal Concept Analysis is applied to a dataset describing the traits of rodents, with the goal of identifying zoonotic disease carriers,or those species carrying infections that can spillover to cause h…

From eye to AI: studying rodent social behavior in the era of machine Learning

2025-08-06 · Giuseppe Chindemi, Camilla Bellone, Benoit Girard arxiv

The study of rodent social behavior has shifted in the last years from relying on direct human observation to more nuanced approaches integrating computational methods in artificial intelligence (AI) and machine learning…