paper-with-me

Papers

Emphasizing Complementary Samples for Non-literal Cross-modal Retrieval

2022-06-25 · IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) 2022 6 · Christopher Thomas, Adriana Kovashka

Existing cross-modal retrieval methods assume a straightforward relationship where images and text contain portrayals or mentions of the same objects. In contrast, real-world image-text pairs (e.g. an image and its caption in a news article) often feature more complex relations. Importantly, not all image-text pairs have the same relationship: in some pairs, image and text may be more closely aligned, while others are more loosely aligned hence complementary. In order to ensure the model learns a semantically robust space which captures nuanced relationships, care must be taken that loosely-aligned image-text pairs have a strong enough impact on learning. In this paper, we propose a novel approach to prioritize loosely-aligned samples. Unlike prior sample weighting methods, ours relies on estimating to what extent semantic similarity is preserved in the separate channels (images/text) in the learned multimodal space. In particular, the image-text pair weights in the retrieval loss focus learning towards samples from diverse or discrepant neighborhoods: samples where images or text that were close in a semantic space, are distant in the crossmodal space (diversity), or where neighbor relations are asymmetric (discrepancy). Experiments on three challenging datasets exhibiting abstract image-text relations, as well as COCO, demonstrate significant performance gains compared to recent state-of-the-art models and sample weighting approaches.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalRetrievalSemantic SimilaritySemantic Textual Similarity

Similar Papers 제목 키워드 기반

Vision, Deduction and Alignment: An Empirical Study on Multi-modal Knowledge Graph Alignment

2023-02-17 · Yangning Li, Jiaoyan Chen, Yinghui Li, Yuejia Xiang 외

Entity alignment (EA) for knowledge graphs (KGs) plays a critical role in knowledge engineering. Existing EA methods mostly focus on utilizing the graph structures and entity attributes (including literals), but ignore i…

Entity AlignmentKnowledge GraphsMulti-modal Knowledge Graph

Beyond the Literal: Decomposing Pragmatic Intent in Multimodal Meme Understanding

2026-06-02 · Zhengyi Zhao, Shubo Zhang, Zezhong Wang, Luyao Ye 외 arxiv

When asked what a meme or sarcastic post means, Large Vision Language Models (LVLMs) tend to describe what the image shows rather than what the author is trying to communicate. Standard instruction tuning entangles a pos…

Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability

2025-04-09 · Ning li, Jingran Zhang, Justin Cui

OpenAI's multimodal GPT-4o has demonstrated remarkable capabilities in image generation and editing, yet its ability to achieve world knowledge-informed semantic synthesis--seamlessly integrating domain knowledge, contex…

Image Generationmultimodal generationWorld Knowledge

Learning Long- and Short-Term User Literal-Preference with Multimodal Hierarchical Transformer Network for Personalized Image Caption

2020-02-04 · AAAI Conference on Artificial Intelligence (AAAI 2020) 2020 2 · Wei Zhang, Yue Ying, Pan Lu, Hongyuan Zha

Personalized image caption, a natural extension of the standard image caption task, requires to generate brief image descriptions tailored for users’ writing style and traits, and is more practical to meet users’ real de…

Image Captioning

Ad Lingua: Text Classification Improves Symbolism Prediction in Image Advertisements

2020-12-01 · COLING 2020 8 · Andrey Savchenko, Anton Alekseev, Sejeong Kwon, Elena Tutubalina 외

Understanding image advertisements is a challenging task, often requiring non-literal interpretation. We argue that standard image-based predictions are insufficient for symbolism prediction. Following the intuition that…

Language ModelingLanguage Modellingobject-detectionObject Detection+4