paper-with-me

홈 › Papers

BiCLIP: Bidirectional and Consistent Language-Image Processing for Robust Medical Image Segmentation

2026-02-25 · Saivan Talaei, Fatemeh Daneshfar, Abdulhady Abas Abdullah, Mustaqeem Khan arxiv

Medical image segmentation is a cornerstone of computer-assisted diagnosis and treatment planning. While recent multimodal vision-language models have shown promise in enhancing semantic understanding through textual descriptions, their resilience in "in-the-wild" clinical settings-characterized by scarce annotations and hardware-induced image degradations-remains under-explored. We introduce BiCLIP (Bidirectional and Consistent Language-Image Processing), a framework engineered to bolster robustness in medical segmentation. BiCLIP features a bidirectional multimodal fusion mechanism that enables visual features to iteratively refine textual representations, ensuring superior semantic alignment. To further stabilize learning, we implement an augmentation consistency objective that regularizes intermediate representations against perturbed input views. Evaluation on the QaTa-COV19 and MosMedData+ benchmarks demonstrates that BiCLIP consistently surpasses state-of-the-art image-only and multimodal baselines. Notably, BiCLIP maintains high performance when trained on as little as 1% of labeled data and exhibits significant resistance to clinical artifacts, including motion blur and low-dose CT noise.

📄 PDF Abstract BibTeX arXiv:2603.00156

Code (0)

등록된 구현이 없습니다.

Tasks

Medical Image Segmentation

Similar Papers 제목 키워드 기반

BiCLIP: Domain Canonicalization via Structured Geometric Transformation

2026-03-09 · Pranav Mantini, Shishir K. Shah arxiv

Recent advances in vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities, yet adapting these models to specialized domains remains a significant challenge. Building on recent theoretical insig…

Domain Adaptation

Mamba in Speech: Towards an Alternative to Self-Attention

2024-05-21 · Xiangyu Zhang, Qiquan Zhang, Hexin Liu, Tianyi Xiao 외

Transformer and its derivatives have achieved success in diverse tasks across computer vision, natural language processing, and speech processing. To reduce the complexity of computations within the multi-head self-atten…

MambaSpeech Enhancementspeech-recognitionSpeech Recognition+1

BRIDLE: Generalized Self-supervised Learning with Quantization

2025-02-04 · Hoang M. Nguyen, Satya N. Shukla, Qiang Zhang, Hanchao Yu 외

Self-supervised learning has been a powerful approach for learning meaningful representations from unlabeled data across various domains, reducing the reliance on large labeled datasets. Inspired by BERT's success in cap…

image-classificationImage ClassificationQuantizationSelf-Supervised Learning+1

Multi-Level Bidirectional Biomimetic Learning for EEG-Based Visual Decoding

2026-05-06 · Jingtao Liu, Peiliang Gong, Chuhang Zheng, Yiheng Liu 외 arxiv

EEG-based visual neural decoding aims to align neural responses with visual stimuli for tasks such as image retrieval. However, limited paired data and a fundamental mismatch between high-fidelity digital images and biol…

Representation LearningContrastive LearningImage Retrieval

GANBERT: Generative Adversarial Networks with Bidirectional Encoder Representations from Transformers for MRI to PET synthesis

2020-08-10 · Hoo-chang Shin, Alvin Ihsani, Swetha Mandava, Sharath Turuvekere Sreenivas 외

Synthesizing medical images, such as PET, is a challenging task due to the fact that the intensity range is much wider and denser than those in photographs and digital renderings and are often heavily biased toward zero.…

Sentence