paper-with-me

Papers

Benchmarking performance, explainability, and evaluation strategies of vision-language models for surgery: Challenges and opportunities

2025-05-16 · Jiajun Cheng, Xianwu Zhao, Shan Lin

Minimally invasive surgery (MIS) presents significant visual and technical challenges, including surgical instrument classification and understanding surgical action involving instruments, verbs, and anatomical targets. While many machine learning-based methods have been developed for surgical understanding, they typically rely on procedure- and task-specific models trained on small, manually annotated datasets. In contrast, the recent success of vision-language models (VLMs) trained on large volumes of raw image-text pairs has demonstrated strong adaptability to diverse visual data and a range of downstream tasks. This opens meaningful research questions: how well do these general-purpose VLMs perform in the surgical domain? In this work, we explore those questions by benchmarking several VLMs across diverse surgical datasets, including general laparoscopic procedures and endoscopic submucosal dissection, to assess their current capabilities and limitations. Our benchmark reveals key gaps in the models' ability to consistently link language to the correct regions in surgical scenes.

📄 PDF Abstract BibTeX arXiv:2505.10764

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models

2025-05-02 · Mahdi Dhaini, Kafaite Zahra Hussain, Efstratios Zaradoukas, Gjergji Kasneci

As Natural Language Processing (NLP) models continue to evolve and become integral to high-stakes applications, ensuring their interpretability remains a critical challenge. Given the growing variety of explainability me…

Benchmarking

GNNX-BENCH: Unravelling the Utility of Perturbation-based GNN Explainers through In-depth Benchmarking

2023-10-03 · Mert Kosan, Samidha Verma, Burouj Armgaan, Khushbu Pahwa 외

Numerous explainability methods have been proposed to shed light on the inner workings of GNNs. Despite the inclusion of empirical evaluations in all the proposed algorithms, the interrogative aspects of these evaluation…

Benchmarkingcounterfactual

FairX: A comprehensive benchmarking tool for model analysis using fairness, utility, and explainability

2024-06-20 · Md Fahim Sikder, Resmi Ramachandranpillai, Daniel de Leng, Fredrik Heintz

We present FairX, an open-source Python-based benchmarking tool designed for the comprehensive analysis of models under the umbrella of fairness, utility, and eXplainability (XAI). FairX enables users to train benchmarki…

BenchmarkingFairness

Dermatological Diagnosis Explainability Benchmark for Convolutional Neural Networks

2023-02-23 · Raluca Jalaboi, Ole Winther, Alfiia Galimzianova

In recent years, large strides have been taken in developing machine learning methods for dermatological applications, supported in part by the success of deep learning (DL). To date, diagnosing diseases from images is o…

BenchmarkingMedical Diagnosis

Synthetic Benchmarks for Scientific Research in Explainable Machine Learning

2021-06-23 · Yang Liu, Sujay Khandagale, Colin White, Willie Neiswanger

As machine learning models grow more complex and their applications become more high-stakes, tools for explaining model predictions have become increasingly important. This has spurred a flurry of research in model expla…

BenchmarkingBIG-bench Machine LearningExplainable Artificial Intelligence (XAI)