paper-with-me

Papers

Multi-Modality Cross Attention Network for Image and Sentence Matching

2020-06-01 · CVPR 2020 6 · Xi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang, Feng Wu

The key of image and sentence matching is to accurately measure the visual-semantic similarity between an image and a sentence. However, most existing methods make use of only the intra-modality relationship within each modality or the inter-modality relationship between image regions and sentence words for the cross-modal matching task. Different from them, in this work, we propose a novel MultiModality Cross Attention (MMCA) Network for image and sentence matching by jointly modeling the intra-modality and inter-modality relationships of image regions and sentence words in a unified deep model. In the proposed MMCA, we design a novel cross-attention mechanism, which is able to exploit not only the intra-modality relationship within each modality, but also the inter-modality relationship between image regions and sentence words to complement and enhance each other for image and sentence matching. Extensive experimental results on two standard benchmarks including Flickr30K and MS-COCO demonstrate that the proposed model performs favorably against state-of-the-art image and sentence matching methods.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SimilaritySemantic Textual SimilaritySentence

Similar Papers 제목 키워드 기반

Modality-specific Cross-modal Similarity Measurement with Recurrent Attention Network

2017-08-16 · Yuxin Peng, Jinwei Qi, Yuxin Yuan

Nowadays, cross-modal retrieval plays an indispensable role to flexibly find information across different modalities of data. Effectively measuring the similarity between different modalities of data is the key of cross-…

Cross-Modal RetrievalRetrievalSentence

ACMM: Aligned Cross-Modal Memory for Few-Shot Image and Sentence Matching

2019-10-01 · ICCV 2019 10 · Yan Huang, Liang Wang

Image and sentence matching has drawn much attention recently, but due to the lack of sufficient pairwise data for training, most previous methods still cannot well associate those challenging pairs of images and sentenc…

cross-modal alignmentSentence

Exploiting Semantic Embedding and Visual Feature for Facial Action Unit Detection

2021-06-19 · CVPR 2021 1 · Huiyuan Yang, Lijun Yin, Yi Zhou, Jiuxiang Gu

Recent study on detecting facial action units (AU) has utilized auxiliary information (i.e., facial landmarks, relationship among AUs and expressions, web facial images, etc.), in order to improve the AU detection pe…

Action Unit DetectionFacial Action Unit DetectionSentence

CMA-CLIP: Cross-Modality Attention CLIP for Image-Text Classification

2021-12-07 · Huidong Liu, Shaoyuan Xu, Jinmiao Fu, Yang Liu 외

Modern Web systems such as social media and e-commerce contain rich contents expressed in images and text. Leveraging information from multi-modalities can improve the performance of machine learning tasks such as classi…

AttributeImage-text ClassificationMultimodal Text and Image Classificationtext-classification+1

Self-Supervised Cross-Modal Text-Image Time Series Retrieval in Remote Sensing

2025-01-31 · Genc Hoxha, Olivér Angyal, Begüm Demir

The development of image time series retrieval (ITSR) methods is a growing research interest in remote sensing (RS). Given a user-defined image time series (i.e., the query time series), the ITSR methods search and retri…

RetrievalTime Series