paper-with-me

홈 › Papers

Scene-adaptive and Region-aware Multi-modal Prompt for Open Vocabulary Object Detection

2024-01-01 · CVPR 2024 1 · Xiaowei Zhao, Xianglong Liu, Duorui Wang, Yajun Gao, Zhide Liu

Open Vocabulary Object Detection (OVD) aims to detect objects from novel classes described by text inputs based on the generalization ability of trained classes. Existing methods mainly focus on transferring knowledge from large Vision and Language models (VLM) to detectors through knowledge distillation. However these approaches show weak ability in adapting to diverse classes and aligning between the image-level pre-training and region-level detection thereby impeding effective knowledge transfer. Motivated by the prompt tuning we propose scene-adaptive and region-aware multi-modal prompts to address these issues by effectively adapting class-aware knowledge from VLM to the detector at the region level. Specifically to enhance the adaptability to diverse classes we design a scene-adaptive prompt generator from a scene perspective to consider both the commonality and diversity of the class distributions and formulate a novel selection mechanism to facilitate the acquisition of common knowledge across all classes and specific insights relevant to each scene. Meanwhile to bridge the gap between the pre-trained model and the detector we present a region-aware multi-modal alignment module which employs the region prompt to incorporate the positional information for feature distillation and integrates textual prompts to align visual and linguistic representations. Extensive experimental results demonstrate that the proposed method significantly outperforms the state-of-the-art models on the OV-COCO and OV-LVIS datasets surpassing the current method by 3.0% mAP and 4.6% APr .

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge Distillationobject-detectionObject DetectionOpen-vocabulary object detectionOpen Vocabulary Object DetectionTransfer Learning

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
Focus 설명 없음

Similar Papers 제목 키워드 기반

Modality-Aware Infrared and Visible Image Fusion with Target-Aware Supervision

2025-09-14 · Tianyao Sun, Dawei Xiang, Tianqi Ding, Xiang Fang 외 arxiv

Infrared and visible image fusion (IVIF) is a fundamental task in multi-modal perception that aims to integrate complementary structural and textural cues from different spectral domains. In this paper, we propose Fusion…

Scene UnderstandingObject Detection

Sub-Region-Aware Modality Fusion and Adaptive Prompting for Multi-Modal Brain Tumor Segmentation

2026-01-22 · Shadi Alijani, Fereshteh Aghaee Meibodi, Homayoun Najjaran arxiv

The successful adaptation of foundation models to multi-modal medical imaging is a critical yet unresolved challenge. Existing models often struggle to effectively fuse information from multiple sources and adapt to the …

Brain Tumor SegmentationPrompt Engineering

SOAP: Vision-Centric 3D Semantic Scene Completion with Scene-Adaptive Decoder and Occluded Region-Aware View Projection

2025-01-01 · CVPR 2025 1 · Hyo-Jun Lee, Yeong Jun Koh, HanUl Kim, Hyunseop Kim 외

Existing view transformations in vision-centric 3D Semantic Scene Completion (SSC) inevitably experience erroneous feature duplication in the reconstructed voxel space due to occlusions, leading to a dilution of info…

3D Semantic Scene CompletionDecoder

RFNet: Region-Aware Fusion Network for Incomplete Multi-Modal Brain Tumor Segmentation

2021-01-01 · ICCV 2021 10 · Yuhang Ding, Xin Yu, Yi Yang

Most existing brain tumor segmentation methods usually exploit multi-modal magnetic resonance imaging (MRI) images to achieve high segmentation performance. However, the problem of missing certain modality images oft…

Brain Tumor SegmentationSegmentationSemantic SegmentationTumor Segmentation

Adaptive and Azimuth-Aware Fusion Network of Multimodal Local Features for 3D Object Detection

2019-10-10 · Yonglin Tian, Kunfeng Wang, Yuang Wang, Yulin Tian 외

This paper focuses on the construction of stronger local features and the effective fusion of image and LiDAR data. We adopt different modalities of LiDAR data to generate richer features and present an adaptive and azim…

3D Object Detectionobject-detectionObject DetectionRegion Proposal