paper-with-me

홈 › Papers

PTCMIL: Multiple Instance Learning via Prompt Token Clustering for Whole Slide Image Analysis

2025-07-24 · Beidi Zhao, SangMook Kim, Hao Chen, Chen Zhou, Zu-hua Gao, Gang Wang, Xiaoxiao Li arxiv

Multiple Instance Learning (MIL) has advanced WSI analysis but struggles with the complexity and heterogeneity of WSIs. Existing MIL methods face challenges in aggregating diverse patch information into robust WSI representations. While ViTs and clustering-based approaches show promise, they are computationally intensive and fail to capture task-specific and slide-specific variability. To address these limitations, we propose PTCMIL, a novel Prompt Token Clustering-based ViT for MIL aggregation. By introducing learnable prompt tokens into the ViT backbone, PTCMIL unifies clustering and prediction tasks in an end-to-end manner. It dynamically aligns clustering with downstream tasks, using projection-based clustering tailored to each WSI, reducing complexity while preserving patch heterogeneity. Through token merging and prototype-based pooling, PTCMIL efficiently captures task-relevant patterns. Extensive experiments on eight datasets demonstrate its superior performance in classification and survival analysis tasks, outperforming state-of-the-art methods. Systematic ablation studies confirm its robustness and strong interpretability. The code is released at https://github.com/ubc-tea/PTCMIL.

📄 PDF Abstract BibTeX arXiv:2507.18848

Code (0)

등록된 구현이 없습니다.

Tasks

Multiple Instance Learning

Similar Papers 제목 키워드 기반

Insight Any Instance: Promptable Instance Segmentation for Remote Sensing Images

2024-09-11 · Xuexue Li

Instance segmentation of remote sensing images (RSIs) is an essential task for a wide range of applications such as land planning and intelligent transport. Instance segmentation of RSIs is constantly plagued by the unba…

Instance SegmentationPrompt LearningSegmentationSemantic Segmentation

TokenCompose: Text-to-Image Diffusion with Token-level Supervision

2023-12-06 · CVPR 2024 1 · ZiRui Wang, Zhizhou Sha, Zheng Ding, Yilin Wang 외

We present TokenCompose, a Latent Diffusion Model for text-to-image generation that achieves enhanced consistency between user-specified text prompts and model-generated images. Despite its tremendous success, the standa…

DenoisingImage GenerationObjectText to Image Generation+1

Hierarchical Prompt Learning for Image- and Text-Based Person Re-Identification

2025-11-17 · Linhan Zhou, Shuang Li, Neng Dong, Yonghang Tai 외 arxiv

Person re-identification (ReID) aims to retrieve target pedestrian images given either visual queries (image-to-image, I2I) or textual descriptions (text-to-image, T2I). Although both tasks share a common retrieval objec…

Person Re-Identification

Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens

2024-05-22 · Qihang Fan, Huaibo Huang, Mingrui Chen, Ran He

The Vision Transformer (ViT) has gained prominence for its superior relational modeling prowess. However, its global attention mechanism's quadratic complexity poses substantial computational burdens. A common remedy spa…

Clusteringimage-classificationImage ClassificationInstance Segmentation+3

Hierarchical and Step-Layer-Wise Tuning of Attention Specialty for Multi-Instance Synthesis in Diffusion Transformers

2025-04-14 · ChunYang Zhang, Zhenhong Sun, Zhicheng Zhang, Junyan Wang 외

Text-to-image (T2I) generation models often struggle with multi-instance synthesis (MIS), where they must accurately depict multiple distinct instances in a single image based on complex prompts detailing individual feat…

AttributeLayout Generation