paper-with-me

홈 › Papers

A Review of 3D Object Detection with Vision-Language Models

2025-04-25 · Ranjan Sapkota, Konstantinos I Roumeliotis, Rahul Harsha Cheppally, Marco Flores Calero, Manoj Karkee

This review provides a systematic analysis of comprehensive survey of 3D object detection with vision-language models(VLMs) , a rapidly advancing area at the intersection of 3D vision and multimodal AI. By examining over 100 research papers, we provide the first systematic analysis dedicated to 3D object detection with vision-language models. We begin by outlining the unique challenges of 3D object detection with vision-language models, emphasizing differences from 2D detection in spatial reasoning and data complexity. Traditional approaches using point clouds and voxel grids are compared to modern vision-language frameworks like CLIP and 3D LLMs, which enable open-vocabulary detection and zero-shot generalization. We review key architectures, pretraining strategies, and prompt engineering methods that align textual and 3D features for effective 3D object detection with vision-language models. Visualization examples and evaluation benchmarks are discussed to illustrate performance and behavior. Finally, we highlight current challenges, such as limited 3D-language datasets and computational demands, and propose future research directions to advance 3D object detection with vision-language models. >Object Detection, Vision-Language Models, Agents, VLMs, LLMs, AI

📄 PDF Abstract BibTeX arXiv:2504.18738

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object DetectionObjectobject-detectionObject DetectionPrompt EngineeringSpatial ReasoningZero-shot Generalization

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Object Detection with Multimodal Large Vision-Language Models: An In-depth Review

2025-08-25 · Ranjan Sapkota, Manoj Karkee arxiv

The fusion of language and vision in large vision-language models (LVLMs) has revolutionized deep learning-based object detection by enhancing adaptability, contextual reasoning, and generalization beyond traditional arc…

Object Detection

Remote Sensing Object Detection Meets Deep Learning: A Meta-review of Challenges and Advances

2023-09-13 · Xiangrong Zhang, Tianyang Zhang, Guanchun Wang, Peng Zhu 외

Remote sensing object detection (RSOD), one of the most fundamental and challenging tasks in the remote sensing field, has received longstanding attention. In recent years, deep learning techniques have demonstrated robu…

Objectobject-detectionObject Detection

Vision-Language Model for Object Detection and Segmentation: A Review and Evaluation

2025-04-13 · Yongchao Feng, Yajie Liu, Shuai Yang, Wenrui Cai 외

Vision-Language Model (VLM) have gained widespread adoption in Open-Vocabulary (OV) object detection and segmentation tasks. Despite they have shown promise on OV-related tasks, their effectiveness in conventional vision…

Domain AdaptationLanguage ModelingLanguage Modellingobject-detection+1

Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing

2024-09-26 · Huthaifa I. Ashqar, Ahmed Jaber, Taqwa I. Alhadidi, Mohammed Elhenawy

This study aims to comprehensively review and empirically evaluate the application of multimodal large language models (MLLMs) and Large Vision Models (VLMs) in object detection for transportation systems. In the first f…

Event DetectionObjectobject-detectionObject Detection+1

Recent Advances of Generic Object Detection with Deep Learning: A Review

2020-12-19 · Xin Li, YingYing Li, Shushu Li

Object detection is an important and challenging problem in computer vision. It has been widely applied in many vision tasks, such as object tracking, image segmentation, action recognition, etc. With the rapid developme…

Action RecognitionData AugmentationDeep LearningImage Segmentation+5