Papers Vision-Language Segmentation
“Vision-Language Segmentation” 태그가 달린 논문 9편 · 필터 해제
HalluSegBench: Counterfactual Visual Reasoning for Segmentation Hallucination Evaluation
Recent progress in vision-language segmentation has significantly advanced grounded visual understanding. However, these models often exhibit hallucinations by producing segmentation masks for objects not grounded in the…
counterfactualCounterfactual ReasoningHallucinationHallucination Evaluation+4RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
Multi-modal Large Language Models (MLLMs) have demonstrated remarkable reasoning capability while lack explicit mechanisms for visual grounding and segmentation, creating a gap between cognitive reasoning and visual perc…
Multimodal ReasoningReasoning SegmentationSegmentationVision-Language Segmentation+2Adversarial Robustness Analysis of Vision-Language Models in Medical Image Segmentation
Adversarial attacks have been fairly explored for computer vision and vision-language models. However, the avenue of adversarial attack for the vision language segmentation models (VLSMs) is still under-explored, especia…
Adversarial AttackAdversarial RobustnessImage SegmentationMedical Image Analysis+3TuneVLSeg: Prompt Tuning Benchmark for Vision-Language Segmentation Models
Vision-Language Models (VLMs) have shown impressive performance in vision tasks, but adapting them to new domains often requires expensive fine-tuning. Prompt tuning techniques, including textual, visual, and multimodal …
BenchmarkingSegmentationVision-Language SegmentationVisual Prompt TuningVLSM-Adapter: Finetuning Vision-Language Segmentation Efficiently with Lightweight Blocks
Foundation Vision-Language Models (VLMs) trained using large-scale open-domain images and text pairs have recently been adapted to develop Vision-Language Segmentation Models (VLSMs) that allow providing text prompts dur…
Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation+1Strong but simple: A Baseline for Domain Generalized Dense Perception by CLIP-based Transfer Learning
Domain generalization (DG) remains a significant challenge for perception based on deep neural networks (DNNs), where domain shifts occur due to synthetic data, lighting, weather, or location changes. Vision-language mod…
Domain Generalizationobject-detectionObject DetectionRobust Object Detection+4Interpretable Diffusion via Information Decomposition
Denoising diffusion models enable conditional generation and density modeling of complex relationships like images and text. However, the nature of the learned relationships is opaque making it difficult to understand pr…
Image GenerationVision-Language SegmentationSynthetic Boost: Leveraging Synthetic Data for Enhanced Vision-Language Segmentation in Echocardiography
Accurate segmentation is essential for echocardiography-based assessment of cardiovascular diseases (CVDs). However, the variability among sonographers and the inherent challenges of ultrasound images hinder precise segm…
SegmentationVision-Language SegmentationExploring Transfer Learning in Medical Image Segmentation using Vision-Language Models
Medical image segmentation allows quantifying target structure size and shape, aiding in disease diagnosis, prognosis, surgery planning, and comprehension.Building upon recent advancements in foundation Vision-Language M…
Image SegmentationMedical Image SegmentationPrognosisSegmentation+3