Image Retrieval
56개 벤치마크 · 논문 2,507편 · 이 태스크의 논문 보기 →
Benchmarks
ROxford (Hard)
ROxford (Medium)
RParis (Hard)
RParis (Medium)
Fashion IQ
CIRR
Flickr30K 1K test
SOP
Flickr30k-CN
Oxf5k
iNaturalist
COCO-CN
Flickr30k
MUGE Retrieval
Oxf105k
CARS196
CUB-200-2011
In-Shop
Par106k
Par6k
AmsterTime
ConQA Conceptual
ConQA Descriptive
PhotoChat
DeepPatent
24/7 Tokyo
Exact Street2Shop
LaSCo
MSCOCO
AIC-ICC
CBVS
INRIA Holidays
Oxford5k
Paris6k
WIT
street2shop - topwear
CIFAR-10
COFAR
DeepFashion
FETA Car-Manuals
FooDI-ML (Global)
FooDI-ML (Spain)
ICFG-PEDES
INSTRE
ImageCoDe
Localized Narratives
NUS-WIDE
PKU SketchRe-ID Dataset
PKU-Reid
RUC-CAS-WenLan
Most implemented
Emerging Properties in Self-Supervised Vision Transformers
DINOv2: Learning Robust Visual Features without Supervision
VGGFace2: A dataset for recognising faces across pose and age
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Circle Loss: A Unified Perspective of Pair Similarity Optimization
Fine-tuning CNN Image Retrieval with No Human Annotation
Papers
PailitaoGR: Latent Think-with-Images for Generative Image Retrieval
Generative retrieval has demonstrated strong performance by directly generating product semantic identifiers (SIDs). Extending this paradigm to image search, however, is nontrivial because real-world query images contain…
Image RetrievalWeaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching
Image retrieval has traditionally been formulated as a point-wise matching problem, where each candidate image is scored in isolation. However, this atomic paradigm fails to capture the complexity of human search intent …
Image RetrievalMulVec: Fine-Grained Role-Aware Matching for Training-Free Zero-Shot Composed Image Retrieval
Training-free zero-shot composed image retrieval finds a target image in a gallery from a reference image and a text edit without learning from task-specific image triplets. Existing methods typically describe the target…
Image RetrievalEviRank: Structured Relevance Evidence for Multimodal Image Re-ranking
Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. Yet existing re-rankers either compress such mul…
Image RetrievalSCORE: Subject Coordinate Recovery for Label-Free Cross-Subject EEG-to-Image Retrieval
Accurate visual decoding can reveal how the brain represents visual information and recover perceived content from neural signals such as electroencephalography (EEG), with potential for neural communication. However, cu…
Image RetrievalBeyond Trial Averaging: Anchoring Neural and Visual Representations for Few-Repetition Brain-to-Image Retrieval
Decoding visual information from brain signals probes neural representations and enables neuro-rehabilitation and dream decoding. Recent brain-to-image retrieval approaches have achieved promising performance, typically …
Image Retrieval