paper-with-me

Cross-Modal Retrieval

13개 벤치마크 · 논문 676편 · 이 태스크의 논문 보기 →

Benchmarks

COCO 2014

결과 108개

Flickr30k

결과 83개

RSICD

결과 31개

RSITMD

결과 31개

Recipe1M

결과 28개

ChEBI-20

결과 27개

MSCOCO-1k

결과 6개

Recipe1M+

결과 6개

SoundingEarth

결과 6개

CUHK-PEDES

결과 3개

Flickr-8k

결과 3개

MS-COCO-2014

결과 3개

MSCOCO

결과 3개

Most implemented

Stacked Capsule Autoencoders

2019-06-17 · 구현 11개

Rescaling Egocentric Vision

2020-06-23 · 구현 7개

Papers

Hub-Spectral Activation of Latent Multimodal Knowledge

2026-09-15 · Ying Guo, Haidong Chen, Linrui Xu, Xiaohao Liu 외 arxiv

Multimodal representation learning seeks shared representations for cross-modal retrieval and knowledge transfer. Hub-based binding reduces pairwise supervision costs, but separate hub connections cannot guarantee reliab…

Representation LearningCross-Modal Retrieval

FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation

2026-09-15 · Guangyu Sun, Shlok Kumar Mishra, Wentao Bao, Robert Zhenheng Yang 외 arxiv

Traditional multimodal representation learning and generation are two stages: a contrastive or self-supervised visual encoder is trained first, followed by a separate downstream generative model. This setup bottlenecks g…

Representation LearningCross-Modal RetrievalImage Captioning

Abstract4D: A Large-Scale Dataset and Framework for Understanding the Visual Language of Abstract Art

2026-08-28 · Haowei Zhang, Yuanpei Zhao, Ji-Zhe Zhou, Mao Li arxiv

Artificial intelligence can classify artistic styles and synthesize images, but it still lacks a model of the visual language that gives art meaning. Abstract painting minimizes object semantics and foregrounds structura…

Text-to-Image GenerationCross-Modal Retrieval

AutoResearch: Insight In, Hallucination Out

2026-08-18 · Yiming Ren, Xiang Liu, Qumeng Sun, Xiao Zhang 외 arxiv

Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two…

Cross-Modal Retrieval

SCALPEL: Semantic Cross-modal Alignment via LLM-Powered Encoder Learning for Medical Vision-Language Representation

2026-07-29 · Yunzhan Fu, Enyu Bao, Xiangyu Shen, Yihao Wu 외 arxiv

Vision-language pre-training (VLP) serves as a cornerstone for medical multimodal representation learning. However, existing medical VLP frameworks are often constrained by the limited context windows and shallow represe…

Visual Question AnsweringRepresentation LearningCross-Modal Retrieval

Not All Patches are Equal: Sampling Matters for Visible-Infrared Pre-Training

2026-07-22 · Qiwei Ma, Bin Deng, Junjie Zhu, Qiangjuan Huang 외 arxiv

Visible-infrared (VIS-IR) alignment is a key pre-training task for robust multi-sensor perception. Most existing methods use uniform patch-wise contrastive learning, but this can be unreliable in VIS-IR data because imag…

Representation LearningCross-Modal RetrievalSemantic SegmentationContrastive Learning

전체 676편 보기 →