paper-with-me

홈 › Papers

A Large-scale Medical Visual Task Adaptation Benchmark

2024-04-19 · Shentong Mo, Xufang Luo, Yansen Wang, Dongsheng Li

Visual task adaptation has been demonstrated to be effective in adapting pre-trained Vision Transformers (ViTs) to general downstream visual tasks using specialized learnable layers or tokens. However, there is yet a large-scale benchmark to fully explore the effect of visual task adaptation on the realistic and important medical domain, particularly across diverse medical visual modalities, such as color images, X-ray, and CT. To close this gap, we present Med-VTAB, a large-scale Medical Visual Task Adaptation Benchmark consisting of 1.68 million medical images for diverse organs, modalities, and adaptation approaches. Based on Med-VTAB, we explore the scaling law of medical prompt tuning concerning tunable parameters and the generalizability of medical visual adaptation using non-medical/medical pre-train weights. Besides, we study the impact of patient ID out-of-distribution on medical visual adaptation, which is a real and challenging scenario. Furthermore, results from Med-VTAB indicate that a single pre-trained model falls short in medical task adaptation. Therefore, we introduce GMoE-Adapter, a novel method that combines medical and general pre-training weights through a gated mixture-of-experts adapter, achieving state-of-the-art results in medical visual task adaptation.

📄 PDF Abstract BibTeX arXiv:2404.12876

Code (0)

등록된 구현이 없습니다.

Tasks

Mixture-of-Experts

Similar Papers 제목 키워드 기반

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation

2025-02-14 · Tianwei Lin, Wenqiao Zhang, Sijing Li, Yuqian Yuan 외

We present HealthGPT, a powerful Medical Large Vision-Language Model (Med-LVLM) that integrates medical visual comprehension and generation capabilities within a unified autoregressive paradigm. Our bootstrapping philoso…

Language ModelingLanguage ModellingPhilosophy

Adapting Visual-Language Models for Generalizable Anomaly Detection in Medical Images

2024-03-19 · CVPR 2024 1 · Chaoqin Huang, Aofan Jiang, Jinghao Feng, Ya zhang 외

Recent advancements in large-scale visual-language pre-trained models have led to significant progress in zero-/few-shot anomaly detection within natural image domains. However, the substantial domain divergence between …

Anomaly ClassificationAnomaly DetectionAnomaly Segmentation

Can Common VLMs Rival Medical VLMs? Evaluation and Strategic Insights

2025-06-19 · Yuan Zhong, Ruinan Jin, Xiaoxiao Li, Qi Dou

Medical vision-language models (VLMs) leverage large-scale pretraining for diverse imaging tasks but require substantial computational and data resources. Meanwhile, common or general-purpose VLMs (e.g., CLIP, LLaVA), th…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models

2026-06-04 · Ryan D'Cunha, Alejandro Lozano, Xiaoxiao Sun, Daniel Vela Jarquin 외 arxiv

Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling cellular features in microscopy. Realizing this potential, however, re…

Domain GeneralizationObject Detection

GRAFT: Grounded and Efficient Online Reinforcement Adaptation for Fine-Grained Robot Manipulation

2026-08-27 · Yibo Qiu, Haoliang Ye, Shu'ang Sun, Zan Huang 외 arxiv

Pretrained vision-language-action (VLA) policies provide strong priors for robot manipulation, yet adapting them online to fine-grained biomedical tasks remains challenging. Task success often hinges on subtle, view-depe…

Robot ManipulationVisual Grounding