paper-with-me

Papers

Exploring Multi-Modal Contextual Knowledge for Open-Vocabulary Object Detection

2023-08-30 · Yifan Xu, Mengdan Zhang, Xiaoshan Yang, Changsheng Xu

In this paper, we for the first time explore helpful multi-modal contextual knowledge to understand novel categories for open-vocabulary object detection (OVD). The multi-modal contextual knowledge stands for the joint relationship across regions and words. However, it is challenging to incorporate such multi-modal contextual knowledge into OVD. The reason is that previous detection frameworks fail to jointly model multi-modal contextual knowledge, as object detectors only support vision inputs and no caption description is provided at test time. To this end, we propose a multi-modal contextual knowledge distillation framework, MMC-Det, to transfer the learned contextual knowledge from a teacher fusion transformer with diverse multi-modal masked language modeling (D-MLM) to a student detector. The diverse multi-modal masked language modeling is realized by an object divergence constraint upon traditional multi-modal masked language modeling (MLM), in order to extract fine-grained region-level visual contexts, which are vital to object detection. Extensive experiments performed upon various detection datasets show the effectiveness of our multi-modal context learning strategy, where our approach well outperforms the recent state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2308.15846

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationLanguage ModelingLanguage ModellingMasked Language ModelingObjectobject-detectionObject DetectionOpen-vocabulary object detectionOpen Vocabulary Object Detection

Methods 이 논문이 사용한 방법론

fail 설명 없음
Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

CrossMuSim: A Cross-Modal Framework for Music Similarity Retrieval with LLM-Powered Text Description Sourcing and Mining

2025-03-29 · Tristan Tsoi, Jiajun Deng, Yaolong Ju, Benno Weck 외

Music similarity retrieval is fundamental for managing and exploring relevant content from large collections in streaming platforms. This paper presents a novel cross-modal contrastive learning framework that leverages t…

Contrastive Learning

COSINT-Agent: A Knowledge-Driven Multimodal Agent for Chinese Open Source Intelligence

2025-03-05 · Wentao Li, Congcong Wang, Xiaoxiao Cui, Zhi Liu 외

Open Source Intelligence (OSINT) requires the integration and reasoning of diverse multimodal data, presenting significant challenges in deriving actionable insights. Traditional approaches, including multimodal large la…

Multimodal Reasoning

Exploring Large Language Models for Multi-Modal Out-of-Distribution Detection

2023-10-12 · Yi Dai, Hao Lang, Kaisheng Zeng, Fei Huang 외

Out-of-distribution (OOD) detection is essential for reliable and trustworthy machine learning. Recent multi-modal OOD detection leverages textual information from in-distribution (ID) class names for visual OOD detectio…

DescriptiveOut-of-Distribution DetectionOut of Distribution (OOD) DetectionWorld Knowledge

Enhancing Cross-Modal Contextual Congruence for Crowdfunding Success using Knowledge-infused Learning

2024-02-06 · Trilok Padhi, Ugur Kursuncu, Yaman Kumar, Valerie L. Shalin 외

The digital landscape continually evolves with multimodality, enriching the online experience for users. Creators and marketers aim to weave subtle contextual cues from various modalities into congruent content to engage…

Common Sense ReasoningKnowledge GraphsMarketing

Exploring Contextual Representation and Multi-Modality for End-to-End Autonomous Driving

2022-10-13 · Shoaib Azam, Farzeen Munir, Ville Kyrki, Moongu Jeon 외

Learning contextual and spatial environmental representations enhances autonomous vehicle's hazard anticipation and decision-making in complex scenarios. Recent perception systems enhance spatial understanding with senso…

Autonomous DrivingSensor Fusion