paper-with-me

홈 › Papers

EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery

2025-01-20 · Guankun Wang, Long Bai, Junyi Wang, Kun Yuan, Zhen Li, Tianxu Jiang, Xiting He, Jinlin Wu, Zhen Chen, Zhen Lei, Hongbin Liu, Jiazheng Wang, Fan Zhang, Nicolas Padoy, Nassir Navab, Hongliang Ren

Recently, Multimodal Large Language Models (MLLMs) have demonstrated their immense potential in computer-aided diagnosis and decision-making. In the context of robotic-assisted surgery, MLLMs can serve as effective tools for surgical training and guidance. However, there is still a lack of MLLMs specialized for surgical scene understanding in clinical applications. In this work, we introduce EndoChat to address various dialogue paradigms and subtasks in surgical scene understanding that surgeons encounter. To train our EndoChat, we construct the Surg-396K dataset through a novel pipeline that systematically extracts surgical information and generates structured annotations based on collected large-scale endoscopic surgery datasets. Furthermore, we introduce a multi-scale visual token interaction mechanism and a visual contrast-based reasoning mechanism to enhance the model's representation learning and reasoning capabilities. Our model achieves state-of-the-art performance across five dialogue paradigms and eight surgical scene understanding tasks. Additionally, we conduct evaluations with professional surgeons, most of whom provide positive feedback on collaborating with EndoChat. Overall, these results demonstrate that our EndoChat has great potential to significantly advance training and automation in robotic-assisted surgery.

📄 PDF Abstract BibTeX arXiv:2501.11347

Code (1)

gkw0010/endochat 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelRepresentation LearningScene Understanding

Similar Papers 제목 키워드 기반

A Multimodal Approach For Endoscopic VCE Image Classification Using BiomedCLIP-PubMedBERT

2024-10-25 · Nagarajan Ganapathy, Podakanti Satyajith Chary, Teja Venkata Ramana Kumar Pithani, Pavan Kavati 외

This Paper presents an advanced approach for fine-tuning BiomedCLIP PubMedBERT, a multimodal model, to classify abnormalities in Video Capsule Endoscopy (VCE) frames, aiming to enhance diagnostic efficiency in gastrointe…

Diagnosticimage-classificationImage ClassificationLanguage Modeling+1

A comprehensive multimodal dataset and benchmark for ulcerative colitis scoring in endoscopy

2026-03-15 · Noha Ghatwary, Jiangbei Yue, Ahmed Elgendy, Hanna Nagdy 외 arxiv

Ulcerative colitis (UC) is a chronic mucosal inflammatory condition that places patients at increased risk of colorectal cancer. Colonoscopic surveillance remains the gold standard for assessing disease activity, and rep…

Image Captioning

BiliVLA: Scene-Aware Vision-Language-Action Model with Reinforcement Learning for Autonomous Biliary Endoscopic Navigation

2026-06-22 · Jinsong Lin, Chi Kit Ng, Zhiyong Xiong, Zikang Pan 외 arxiv

Endoscopic retrograde cholangiopancreatography (ERCP) demands precise endoscopic navigation and stable biliary cannulation within a narrow monocular field characterized by specular reflections, partial occlusions, and fr…

Reinforcement Learning

Registration-Grounded Spectral Fusion for Unregistered WLI/NBI Endoscopic Lesion Segmentation

2026-07-29 · Pengyu Jie, Wanquan Liu, Rui He, Pengcheng Li 외 arxiv

White-light imaging (WLI) and narrow-band imaging (NBI) provide complementary views of endoscopic lesions, but their paired observations are often spatially misaligned due to viewpoint changes, tissue deformation, and se…

Lesion Segmentation

How can reasoning capability empower the AI copilot robot in endoscopic surgery

2026-05-21 · Guankun Wang, Long Bai, Hongliang Ren arxiv

Reasoning capability has significantly advanced complex logical inference and robotic decision-making in general domains. However, its potential in the Artificial Intelligence (AI) copilot robot-particularly implemented …