paper-with-me

홈 › Papers

Vision-Language Model for Object Detection and Segmentation: A Review and Evaluation

2025-04-13 · Yongchao Feng, Yajie Liu, Shuai Yang, Wenrui Cai, Jinqing Zhang, Qiqi Zhan, Ziyue Huang, Hongxi Yan, Qiao Wan, ChenGuang Liu, Junzhe Wang, Jiahui Lv, Ziqi Liu, Tengyuan Shi, Qingjie Liu, Yunhong Wang

Vision-Language Model (VLM) have gained widespread adoption in Open-Vocabulary (OV) object detection and segmentation tasks. Despite they have shown promise on OV-related tasks, their effectiveness in conventional vision tasks has thus far been unevaluated. In this work, we present the systematic review of VLM-based detection and segmentation, view VLM as the foundational model and conduct comprehensive evaluations across multiple downstream tasks for the first time: 1) The evaluation spans eight detection scenarios (closed-set detection, domain adaptation, crowded objects, etc.) and eight segmentation scenarios (few-shot, open-world, small object, etc.), revealing distinct performance advantages and limitations of various VLM architectures across tasks. 2) As for detection tasks, we evaluate VLMs under three finetuning granularities: \textit{zero prediction}, \textit{visual fine-tuning}, and \textit{text prompt}, and further analyze how different finetuning strategies impact performance under varied task. 3) Based on empirical findings, we provide in-depth analysis of the correlations between task characteristics, model architectures, and training methodologies, offering insights for future VLM design. 4) We believe that this work shall be valuable to the pattern recognition experts working in the fields of computer vision, multimodal learning, and vision foundation models by introducing them to the problem, and familiarizing them with the current status of the progress while providing promising directions for future research. A project associated with this review and evaluation has been created at https://github.com/better-chao/perceptual_abilities_evaluation.

📄 PDF Abstract BibTeX arXiv:2504.09480

Code (1)

better-chao/perceptual_abilities_evaluation 공식 구현 pytorch

Tasks

Domain AdaptationLanguage ModelingLanguage Modellingobject-detectionObject Detection

Similar Papers 제목 키워드 기반

Object Detection with Multimodal Large Vision-Language Models: An In-depth Review

2025-08-25 · Ranjan Sapkota, Manoj Karkee arxiv

The fusion of language and vision in large vision-language models (LVLMs) has revolutionized deep learning-based object detection by enhancing adaptability, contextual reasoning, and generalization beyond traditional arc…

Object Detection

A Review of 3D Object Detection with Vision-Language Models

2025-04-25 · Ranjan Sapkota, Konstantinos I Roumeliotis, Rahul Harsha Cheppally, Marco Flores Calero 외

This review provides a systematic analysis of comprehensive survey of 3D object detection with vision-language models(VLMs) , a rapidly advancing area at the intersection of 3D vision and multimodal AI. By examining over…

3D Object DetectionObjectobject-detectionObject Detection+3

Recent Advances of Generic Object Detection with Deep Learning: A Review

2020-12-19 · Xin Li, YingYing Li, Shushu Li

Object detection is an important and challenging problem in computer vision. It has been widely applied in many vision tasks, such as object tracking, image segmentation, action recognition, etc. With the rapid developme…

Action RecognitionData AugmentationDeep LearningImage Segmentation+5

Vision Transformers: State of the Art and Research Challenges

2022-07-07 · Bo-Kai Ruan, Hong-Han Shuai, Wen-Huang Cheng

Transformers have achieved great success in natural language processing. Due to the powerful capability of self-attention mechanism in transformers, researchers develop the vision transformers for a variety of computer v…

3D ReconstructionImage Segmentationobject-detectionObject Detection+3

Deep learning approaches to surgical video segmentation and object detection: A Scoping Review

2025-02-23 · Devanish N. Kamtam, Joseph B. Shrager, Satya Deepya Malla, Nicole Lin 외

Introduction: Computer vision (CV) has had a transformative impact in biomedical fields such as radiology, dermatology, and pathology. Its real-world adoption in surgical applications, however, remains limited. We review…

object-detectionObject DetectionSegmentationSemantic Segmentation+2