paper-with-me

홈 › Papers

NeuroABench: A Multimodal Evaluation Benchmark for Neurosurgical Anatomy Identification

2025-12-07 · Ziyang Song, Zelin Zang, Xiaofan Ye, Boqiang Xu, Long Bai, Jinlin Wu, Hongliang Ren, Hongbin Liu, Jiebo Luo, Zhen Lei arxiv

Multimodal Large Language Models (MLLMs) have shown significant potential in surgical video understanding. With improved zero-shot performance and more effective human-machine interaction, they provide a strong foundation for advancing surgical education and assistance. However, existing research and datasets primarily focus on understanding surgical procedures and workflows, while paying limited attention to the critical role of anatomical comprehension. In clinical practice, surgeons rely heavily on precise anatomical understanding to interpret, review, and learn from surgical videos. To fill this gap, we introduce the Neurosurgical Anatomy Benchmark (NeuroABench), the first multimodal benchmark explicitly created to evaluate anatomical understanding in the neurosurgical domain. NeuroABench consists of 9 hours of annotated neurosurgical videos covering 89 distinct procedures and is developed using a novel multimodal annotation pipeline with multiple review cycles. The benchmark evaluates the identification of 68 clinical anatomical structures, providing a rigorous and standardized framework for assessing model performance. Experiments on over 10 state-of-the-art MLLMs reveal significant limitations, with the best-performing model achieving only 40.87% accuracy in anatomical identification tasks. To further evaluate the benchmark, we extract a subset of the dataset and conduct an informative test with four neurosurgical trainees. The results show that the best-performing student achieves 56% accuracy, with the lowest scores of 28% and an average score of 46.5%. While the best MLLM performs comparably to the lowest-scoring student, it still lags significantly behind the group's average performance. This comparison underscores both the progress of MLLMs in anatomical understanding and the substantial gap that remains in achieving human-level performance.

📄 PDF Abstract BibTeX arXiv:2512.06921

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Data-Driven Registration and Modeling of Brain Deformation for Image-Guided Neurosurgery: A Systematic Review

2026-02-09 · Tiago Assis, Colin P. Galvin, Joshua P. Castillo, Nazim Haouchine 외 arxiv

Accurate compensation of brain deformation is critical for reliable image-guided neurosurgery. Surgical manipulation and tumor resection induce tissue motion, causing preoperative planning images to become misaligned wit…

Computational EfficiencyImage Registration

Vision-Based Neurosurgical Guidance: Unsupervised Localization and Camera-Pose Prediction

2024-05-15 · Gary Sarwin, Alessandro Carretta, Victor Staartjes, Matteo Zoli 외

Localizing oneself during endoscopic procedures can be problematic due to the lack of distinguishable textures and landmarks, as well as difficulties due to the endoscopic device such as a limited field of view and chall…

AnatomyPose Prediction

Zero-shot generation of synthetic neurosurgical data with large language models

2025-02-13 · Austin A. Barr, Eddie Guo, Emre Sezgin

Clinical data is fundamental to advance neurosurgical research, but access is often constrained by data availability, small sample sizes, privacy regulations, and resource-intensive preprocessing and de-identification pr…

BenchmarkingDe-identificationGenerative Adversarial NetworkLarge Language Model

CrossMed: A Multimodal Cross-Task Benchmark for Compositional Generalization in Medical Imaging

2025-11-14 · Pooja Singh, Siddhant Ujjain, Tapan Kumar Gandhi, Sandeep Kumar arxiv

Recent advances in multimodal large language models have enabled unified processing of visual and textual inputs, offering promising applications in general-purpose medical AI. However, their ability to generalize compos…

Visual Question Answering

6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models

2025-12-03 · Leon Mayer, Piotr Kalinowski, Caroline Ebersbach, Marcel Knopp 외 arxiv

Vision-language models (VLMs) are increasingly integrated into clinical workflows. However, existing benchmarks primarily assess performance on common anatomical presentations and fail to capture the challenges posed by …