paper-with-me

홈 › Papers

BiomedAP: A Vision-Informed Dual-Anchor Framework with Gated Cross-Modal Fusion for Robust Medical Vision-Language Adaptation

2026-05-15 · Huanyang Tong, Kai Liu, Fangjun Kuang, Huiling Chen arxiv

Biomedical Vision--Language Models (VLMs) have shown remarkable promise in few-shot medical diagnosis but face a critical bottleneck: \textit{fragility to prompt variations}.Existing adaptation frameworks typically optimize visual and textual prompts as independent streams, relying on ideal ``Golden Prompts''. In clinical reality, where descriptions are often noisy and heterogeneous, this modality isolation leads to unstable cross-modal alignment. To address this, we propose BiomedAP, a vision-informed dual-anchor framework with gated cross-modal fusion.BiomedAP enforces synergistic alignment through two mechanisms: (1) Gated Cross-Modal Fusion, which enables layer-wise interaction between modalities, acting as a dynamic noise regulator to suppress irrelevant textual cues; and (2) a Dual-Anchor Constraint that regularizes learnable prompts toward stable semantic centroids derived from both expert templates (High Anchors) and few-shot visual prototypes (Low Anchors). Extensive experiments across 11 benchmarks demonstrate that BiomedAP consistently surpasses baselines, achieving competitive few-shot accuracy and markedly enhanced robustness under prompt perturbations. Our code is available at: https://github.com/tongdiedie/BiomedAP. Keywords: Vision-Language Models; Prompt Learning; Parameter-Efficient Fine-Tuning; Few-shot Learning

📄 PDF Abstract BibTeX arXiv:2605.15736

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuningFew-Shot LearningMedical Diagnosis

Similar Papers 제목 키워드 기반

EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation

2026-03-10 · Jiajun Cao, Xiaoan Zhang, Xiaobao Wei, Liyuqiu Huang 외 arxiv

Vision-Language-Action models have shown great promise for autonomous driving, yet they suffer from degraded perception after unfreezing the visual encoder and struggle with accumulated instability in long-term planning.…

Autonomous Driving

The Evaluator Is Part of the Experiment: Measuring Open-Ended LLM Conformity

2026-08-05 · Alicia Guerra, Yibo Hu arxiv

Prior work on LLM conformity largely measures discrete answer flips under verifiable labels. Open-ended revisions require a different measurement strategy because answer quality is graded, latent, and judged imperfectly.…

Dual-Modality Anchor-Guided Filtering for Test-time Prompt Tuning

2026-04-14 · Jungwon Choi, Eunwoo Kim arxiv

Test-Time Prompt Tuning (TPT) adapts vision-language models using augmented views, but its effectiveness is hindered by the challenge of determining which views are beneficial. Standard entropy-based filtering relies on …

ANCHOR: Error-Controlled Adaptive Numerical Correction for Neural Operator Time Marching

2025-12-22 · Rajyasri Roy, Dibyajyoti Nayak, Somdatta Goswami arxiv

Numerical simulation of time-dependent partial differential equations (PDEs) is central to scientific and engineering applications, but high-fidelity solvers are often prohibitively expensive for long-horizon or time-cri…

DREAM: Dual-Standard Semantic Homogeneity with Dynamic Optimization for Graph Learning with Label Noise

2026-01-24 · Yusheng Zhao, Jiaye Xie, Qixin Zhang, Weizhi Zhang 외 arxiv

Graph neural networks (GNNs) have been widely used in various graph machine learning scenarios. Existing literature primarily assumes well-annotated training graphs, while the reliability of labels is not guaranteed in r…

Graph Learning