paper-with-me

Papers

A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports

2026-07-30 · Jia Yu, Yan Zhu, Yili He, Zilong Wang, Xinyang Jiang, Peiyao Fu, Ruijie Yang, Tianyi Chen, Siyuan Li, Zhihua Wang, Fei Wu, Quanlin Li, Xian Yang, Pinghong Zhou, Shuo Wang arxiv

Vision-language models remain underused in colonoscopy despite the rich expert descriptions recorded in routine reports. These reports document lesion appearance, size and location but summarise entire procedures rather than caption individual frames, leaving clinical findings only weakly linked to the corresponding images. Here we develop EndoCLIP, a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs progressively recovered from 280,476 routine colonoscopy records. Across lesion-level image-text retrieval, structured report generation and six multi-centre clinical classification tasks, EndoCLIP outperforms general-purpose and biomedical vision-language encoders in both zero-shot and linear-probe settings. On benign-versus-malignant classification, its linear probe approaches the performance of expert readers in a blinded study involving 12 endoscopists. These results suggest that recovering finding-to-frame correspondence can transform routine documentation into scalable supervision, enabling clinical targets to be specified in language rather than separately annotated for each task.

📄 PDF Abstract BibTeX arXiv:2607.28466

Code (2)

InsomaniacElf/sg-tamil-tts-resources- ★ 1
Tavish9/awesome-daily-AI-arxiv ★ 112

Tasks

Text Retrieval

Similar Papers 제목 키워드 기반

Knowledge Extraction and Distillation from Large-Scale Image-Text Colonoscopy Records Leveraging Large Language and Vision Models

2023-10-17 · Shuo Wang, Yan Zhu, Xiaoyuan Luo, Zhiwei Yang 외

The development of artificial intelligence systems for colonoscopy analysis often necessitates expert-annotated image datasets. However, limitations in dataset size and diversity impede model performance and generalisati…

Diversity

Frontiers in Intelligent Colonoscopy

2024-10-22 · Ge-Peng Ji, Jingyi Liu, Peng Xu, Nick Barnes 외

Colonoscopy is currently one of the most sensitive screening methods for colorectal cancer. This study investigates the frontiers of intelligent colonoscopy techniques and their prospective implications for multimodal me…

Image CaptioningImage ClassificationLanguage Modeling+3

Colon-X: Advancing Intelligent Colonoscopy toward Clinical Reasoning

2025-12-03 · Ge-Peng Ji, Jingyi Liu, Deng-Ping Fan, Huazhu Fu 외 arxiv

In this study, we present Colon-X, an open initiative aimed at advancing multimodal intelligence in colonoscopy. We begin by constructing ColonVQA, the most comprehensive multimodal dataset ever built for colonoscopy, fe…

Visual Question Answering

OpenRC: An Open-Source Robotic Colonoscopy Framework for Multimodal Data Acquisition and Autonomy Research

2026-04-04 · Siddhartha Kapuria, Mohammad Rafiee Javazm, Naruhiko Ikoma, Joga Ivatury 외 arxiv

Colorectal cancer screening critically depends on colonoscopy, yet existing platforms offer limited support for systematically studying the coupled dynamics of operator control, instrument motion, and visual feedback. Th…

Spectral Rectification for Parameter-Efficient Adaptation of Foundation Models in Colonoscopy Depth Estimation

2026-03-16 · Xiaoxian Zhang, Minghai Shi, Lei Li arxiv

Accurate monocular depth estimation is critical in colonoscopy for lesion localization and navigation. Foundation models trained on natural images fail to generalize directly to colonoscopy. We identify the core issue no…

Monocular Depth Estimation