paper-with-me

홈 › Papers

Cross-Modal Hierarchical Modelling for Fine-Grained Sketch Based Image Retrieval

2020-07-29 · Aneeshan Sain, Ayan Kumar Bhunia, Yongxin Yang, Tao Xiang, Yi-Zhe Song

Sketch as an image search query is an ideal alternative to text in capturing the fine-grained visual details. Prior successes on fine-grained sketch-based image retrieval (FG-SBIR) have demonstrated the importance of tackling the unique traits of sketches as opposed to photos, e.g., temporal vs. static, strokes vs. pixels, and abstract vs. pixel-perfect. In this paper, we study a further trait of sketches that has been overlooked to date, that is, they are hierarchical in terms of the levels of detail -- a person typically sketches up to various extents of detail to depict an object. This hierarchical structure is often visually distinct. In this paper, we design a novel network that is capable of cultivating sketch-specific hierarchies and exploiting them to match sketch with photo at corresponding hierarchical levels. In particular, features from a sketch and a photo are enriched using cross-modal co-attention, coupled with hierarchical node fusion at every level to form a better embedding space to conduct retrieval. Experiments on common benchmarks show our method to outperform state-of-the-arts by a significant margin.

📄 PDF Abstract BibTeX arXiv:2007.15103

Code (1)

aneeshan95/Cross-modal_Hierarchy_FGSBIR pytorch

Tasks

Image RetrievalRetrievalSketch-Based Image Retrieval

Similar Papers 제목 키워드 기반

Fine-grained Video-Text Retrieval with Hierarchical Graph Reasoning

2020-03-01 · CVPR 2020 6 · Shizhe Chen, Yida Zhao, Qin Jin, Qi Wu

Cross-modal retrieval between videos and texts has attracted growing attentions due to the rapid emergence of videos on the web. The current dominant approach for this problem is to learn a joint embedding space to measu…

Cross-Modal RetrievalRetrievalText MatchingText Retrieval+1

Achieving Fine-grained Cross-modal Understanding through Brain-inspired Hierarchical Representation Learning

2026-01-04 · Weihang You, Hanqi Jiang, Yi Pan, Junhao Chen 외 arxiv

Understanding neural responses to visual stimuli remains challenging due to the inherent complexity of brain representations and the modality gap between neural data and visual inputs. Existing methods, mainly based on r…

Representation LearningCross-Modal RetrievalContrastive LearningVideo Alignment

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning

2024-12-30 · Peng Jin, Hao Li, Li Yuan, Shuicheng Yan 외

Multimodal representation learning, with contrastive learning, plays an important role in the artificial intelligence domain. As an important subfield, video-language representation learning focuses on learning represent…

Contrastive LearningQuestion AnsweringRepresentation LearningVideo Captioning+2

Tencent Text-Video Retrieval: Hierarchical Cross-Modal Interactions with Multi-Level Representations

2022-04-07 · Jie Jiang, Shaobo Min, Weijie Kong, Dihong Gong 외

Text-Video Retrieval plays an important role in multi-modal understanding and has attracted increasing attention in recent years. Most existing methods focus on constructing contrastive pairs between whole videos and com…

Contrastive LearningDenoisingRetrievalSentence+2

HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual Grounding

2024-04-20 · Linhui Xiao, Xiaoshan Yang, Fang Peng, YaoWei Wang 외

Visual grounding, which aims to ground a visual region via natural language, is a task that heavily relies on cross-modal alignment. Existing works utilized uni-modal pre-trained models to transfer visual or linguistic k…

cross-modal alignmentVisual Grounding