paper-with-me

Papers

Quilt-1M: One Million Image-Text Pairs for Histopathology

2023-06-20 · NeurIPS 2023 11 · Wisdom Oluchi Ikezogwo, Mehmet Saygin Seyfioglu, Fatemeh Ghezloo, Dylan Stefan Chan Geva, Fatwir Sheikh Mohammed, Pavan Kumar Anand, Ranjay Krishna, Linda Shapiro

Recent accelerations in multi-modal applications have been made possible with the plethora of image and text data available online. However, the scarcity of analogous data in the medical field, specifically in histopathology, has slowed comparable progress. To enable similar representation learning for histopathology, we turn to YouTube, an untapped resource of videos, offering $1,087$ hours of valuable educational histopathology videos from expert clinicians. From YouTube, we curate QUILT: a large-scale vision-language dataset consisting of $802, 144$ image and text pairs. QUILT was automatically curated using a mixture of models, including large language models, handcrafted algorithms, human knowledge databases, and automatic speech recognition. In comparison, the most comprehensive datasets curated for histopathology amass only around $200$K samples. We combine QUILT with datasets from other sources, including Twitter, research papers, and the internet in general, to create an even larger dataset: QUILT-1M, with $1$M paired image-text samples, marking it as the largest vision-language histopathology dataset to date. We demonstrate the value of QUILT-1M by fine-tuning a pre-trained CLIP model. Our model outperforms state-of-the-art models on both zero-shot and linear probing tasks for classifying new histopathology images across $13$ diverse patch-level datasets of $8$ different sub-pathologies and cross-modal retrieval tasks.

📄 PDF Abstract BibTeX arXiv:2306.11207

Code (2)

wisdomikezogwo/quilt1m 공식 구현 pytorch
pathmmu-benchmark/pathmmu

Tasks

Automatic Speech RecognitionCross-Modal RetrievalRepresentation Learningspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos

2023-12-07 · CVPR 2024 1 · Mehmet Saygin Seyfioglu, Wisdom O. Ikezogwo, Fatemeh Ghezloo, Ranjay Krishna 외

Diagnosis in histopathology requires a global whole slide images (WSIs) analysis, requiring pathologists to compound evidence from different WSI patches. The gigapixel scale of WSIs poses a challenge for histopathology m…

DiagnosticImage CaptioningVisual Question Answering (VQA)whole slide images

Effortless Vision-Language Model Specialization in Histopathology without Annotation

2025-08-11 · Jingna Qiu, Nishanth Jain, Jonas Ammeling, Marc Aubreville 외 arxiv

Recent advances in Vision-Language Models (VLMs) in histopathology, such as CONCH and QuiltNet, have demonstrated impressive zero-shot classification capabilities across various tasks. However, their general-purpose desi…

Model-based Cleaning of the QUILT-1M Pathology Dataset for Text-Conditional Image Synthesis

2024-04-11 · Marc Aubreville, Jonathan Ganz, Jonas Ammeling, Christopher C. Kaltenecker 외

The QUILT-1M dataset is the first openly available dataset containing images harvested from various online sources. While it provides a huge data variety, the image quality and composition is highly heterogeneous, impact…

Image Generation

Towards a Visual-Language Foundation Model for Computational Pathology

2023-07-24 · Ming Y. Lu, Bowen Chen, Drew F. K. Williamson, Richard J. Chen 외

The accelerated adoption of digital pathology and advances in deep learning have enabled the development of powerful models for various pathology tasks across a diverse array of diseases and patient cohorts. However, mod…

Contrastive Learningimage-classificationImage ClassificationImage to text+3

A Generative Foundation Model for Multimodal Histopathology

2026-04-04 · Jinxi Xiang, Mingjie Li, Siyu Hou, Yijiang Chen 외 arxiv

Accurate diagnosis and treatment of complex diseases require integrating histological, molecular, and clinical data, yet in practice these modalities are often incomplete owing to tissue scarcity, assay cost, and workflo…

Data AugmentationImage Generation