paper-with-me

홈 › Papers

ParNet: Position-aware Aggregated Relation Network for Image-Text matching

2019-06-17 · Yaxian Xia, Lun Huang, Xiao-Yong Wei, Wenmin Wang

Exploring fine-grained relationship between entities(e.g. objects in image or words in sentence) has great contribution to understand multimedia content precisely. Previous attention mechanism employed in image-text matching either takes multiple self attention steps to gather correspondences or uses image objects (or words) as context to infer image-text similarity. However, they only take advantage of semantic information without considering that objects' relative position also contributes to image understanding. To this end, we introduce a novel position-aware relation module to model both the semantic and spatial relationship simultaneously for image-text matching in this paper. Given an image, our method utilizes the location of different objects to capture spatial relationship innovatively. With the combination of semantic and spatial relationship, it's easier to understand the content of different modalities (images and sentences) and capture fine-grained latent correspondences of image-text pairs. Besides, we employ a two-step aggregated relation module to capture interpretable alignment of image-text pairs. The first step, we call it intra-modal relation mechanism, in which we computes responses between different objects in an image or different words in a sentence separately; The second step, we call it inter-modal relation mechanism, in which the query plays a role of textual context to refine the relationship among object proposals in an image. In this way, our position-aware aggregated relation network (ParNet) not only knows which entities are relevant by attending on different objects (words) adaptively, but also adjust the inter-modal correspondence according to the latent alignments according to query's content. Our approach achieves the state-of-the-art results on MS-COCO dataset.

📄 PDF Abstract BibTeX arXiv:1906.06892

Code (0)

등록된 구현이 없습니다.

Tasks

Image-text matchingPositionRelationRelation NetworkSentenceText Matchingtext similarity

Similar Papers 제목 키워드 기반

Infrared Image Deturbulence Restoration Using Degradation Parameter-Assisted Wide & Deep Learning

2023-05-30 · Yi Lu, Yadong Wang, Xingbo Jiang, Xiangzhi Bai

Infrared images captured under turbulent conditions are degraded by complex geometric distortions and blur. We address infrared deturbulence as an image restoration task, proposing DparNet, a parameter-assisted multi-fra…

Deep LearningDenoisingImage DenoisingImage Restoration

Optimal operating MR contrast for brain ventricle parcellation

2023-04-04 · Savannah P. Hays, Lianrui Zuo, Yuli Wang, Mark G. Luciano 외

Development of MR harmonization has enabled different contrast MRIs to be synthesized while preserving the underlying anatomy. In this paper, we use image harmonization to explore the impact of different T1-w MR contrast…

AnatomyImage Harmonization

Learning Spatial Attention for Face Super-Resolution

2020-12-02 · Chaofeng Chen, Dihong Gong, Hao Wang, Zhifeng Li 외

General image super-resolution techniques have difficulties in recovering detailed face structures when applying to low resolution face images. Recent deep learning based methods tailored for face images have achieved im…

Face ParsingImage Super-ResolutionMulti-Task LearningSSIM+1

SPARNet: Continual Test-Time Adaptation via Sample Partitioning Strategy and Anti-Forgetting Regularization

2025-01-01 · Xinru Meng, Han Sun, Jiamei Liu, Ningzhong Liu 외

Test-time Adaptation (TTA) aims to improve model performance when the model encounters domain changes after deployment. The standard TTA mainly considers the case where the target domain is static, while the continual TT…

Test-time Adaptation

SALAD -- Semantics-Aware Logical Anomaly Detection

2025-09-02 · Matic Fučka, Vitjan Zavrtanik, Danijel Skočaj arxiv

Recent surface anomaly detection methods excel at identifying structural anomalies, such as dents and scratches, but struggle with logical anomalies, such as irregular or missing object components. The best-performing lo…

Anomaly Detection