paper-with-me

홈 › Papers

LMM-Det: Make Large Multimodal Models Excel in Object Detection

2025-07-24 · Jincheng Li, Chunyu Xie, Ji Ao, Dawei Leng, Yuhui Yin arxiv

Large multimodal models (LMMs) have garnered wide-spread attention and interest within the artificial intelligence research and industrial communities, owing to their remarkable capability in multimodal understanding, reasoning, and in-context learning, among others. While LMMs have demonstrated promising results in tackling multimodal tasks like image captioning, visual question answering, and visual grounding, the object detection capabilities of LMMs exhibit a significant gap compared to specialist detectors. To bridge the gap, we depart from the conventional methods of integrating heavy detectors with LMMs and propose LMM-Det, a simple yet effective approach that leverages a Large Multimodal Model for vanilla object Detection without relying on specialized detection modules. Specifically, we conduct a comprehensive exploratory analysis when a large multimodal model meets with object detection, revealing that the recall rate degrades significantly compared with specialist detection models. To mitigate this, we propose to increase the recall rate by introducing data distribution adjustment and inference optimization tailored for object detection. We re-organize the instruction conversations to enhance the object detection capabilities of large multimodal models. We claim that a large multimodal model possesses detection capability without any extra detection modules. Extensive experiments support our claim and show the effectiveness of the versatile LMM-Det. The datasets, models, and codes are available at https://github.com/360CVGroup/LMM-Det.

📄 PDF Abstract BibTeX arXiv:2507.18300

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringVisual GroundingObject DetectionImage Captioning

Similar Papers 제목 키워드 기반

ReasonCD: A Multimodal Reasoning Large Model for Implicit Change-of-Interest Semantic Mining

2025-12-22 · Zhenyang Huang, Xiao Yu, Yi Zhang, Decheng Wang 외 arxiv

Remote sensing image change detection is one of the fundamental tasks in remote sensing intelligent interpretation. Its core objective is to identify changes within change regions of interest (CRoI). Current multimodal l…

Multimodal ReasoningChange Detection

ROD-MLLM: Towards More Reliable Object Detection in Multimodal Large Language Models

2025-01-01 · CVPR 2025 1 · Heng Yin, Yuqiang Ren, Ke Yan, Shouhong Ding 외

Multimodal large language models (MLLMs) have demonstrated strong language understanding and generation capabilities, excelling in visual tasks like referring and grounding. However, due to task type limitations and …

Large Language ModelObjectobject-detectionObject Detection

Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object Detection

2025-03-10 · Wentao Wu, Chenglong Li, Xiao Wang, Bin Luo 외

Existing multimodal UAV object detection methods often overlook the impact of semantic gaps between modalities, which makes it difficult to achieve accurate semantic and spatial alignments, limiting detection performance…

Language ModelingLanguage ModellingLarge Language ModelObject+2

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning

2024-11-15 · Jingru Yang, Huan Yu, Yang Jingxin, Chentianye Xu 외

Multimodal Large Language Models (MLLMs) excel at descriptive tasks within images but often struggle with precise object localization, a critical element for reliable visual interpretation. In contrast, traditional objec…

DescriptiveObjectobject-detectionObject Detection+3

EagleVision: Object-level Attribute Multimodal LLM for Remote Sensing

2025-03-30 · Hongxiang Jiang, Jihao Yin, Qixiong Wang, Jiaqi Feng 외

Recent advances in multimodal large language models (MLLMs) have demonstrated impressive results in various visual tasks. However, in remote sensing (RS), high resolution and small proportion of objects pose challenges t…

AttributeDisentanglementObjectobject-detection+1