Cross-Modal Retrieval
13개 벤치마크 · 논문 676편 · 이 태스크의 논문 보기 →
Benchmarks
COCO 2014
Flickr30k
RSICD
RSITMD
Recipe1M
ChEBI-20
MSCOCO-1k
Recipe1M+
SoundingEarth
CUHK-PEDES
Flickr-8k
MS-COCO-2014
MSCOCO
Most implemented
Stacked Capsule Autoencoders
VSE++: Improving Visual-Semantic Embeddings with Hard Negatives
Rescaling Egocentric Vision
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
Stacked Cross Attention for Image-Text Matching
Papers
Hub-Spectral Activation of Latent Multimodal Knowledge
Multimodal representation learning seeks shared representations for cross-modal retrieval and knowledge transfer. Hub-based binding reduces pairwise supervision costs, but separate hub connections cannot guarantee reliab…
Representation LearningCross-Modal RetrievalFLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation
Traditional multimodal representation learning and generation are two stages: a contrastive or self-supervised visual encoder is trained first, followed by a separate downstream generative model. This setup bottlenecks g…
Representation LearningCross-Modal RetrievalImage CaptioningAbstract4D: A Large-Scale Dataset and Framework for Understanding the Visual Language of Abstract Art
Artificial intelligence can classify artistic styles and synthesize images, but it still lacks a model of the visual language that gives art meaning. Abstract painting minimizes object semantics and foregrounds structura…
Text-to-Image GenerationCross-Modal RetrievalAutoResearch: Insight In, Hallucination Out
Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two…
Cross-Modal RetrievalSCALPEL: Semantic Cross-modal Alignment via LLM-Powered Encoder Learning for Medical Vision-Language Representation
Vision-language pre-training (VLP) serves as a cornerstone for medical multimodal representation learning. However, existing medical VLP frameworks are often constrained by the limited context windows and shallow represe…
Visual Question AnsweringRepresentation LearningCross-Modal RetrievalNot All Patches are Equal: Sampling Matters for Visible-Infrared Pre-Training
Visible-infrared (VIS-IR) alignment is a key pre-training task for robust multi-sensor perception. Most existing methods use uniform patch-wise contrastive learning, but this can be unreliable in VIS-IR data because imag…
Representation LearningCross-Modal RetrievalSemantic SegmentationContrastive Learning