paper-with-me

홈 › Papers

FIGROTD: A Friendly-to-Handle Dataset for Image Guided Retrieval with Optional Text

2025-11-27 · Hoang-Bao Le, Allie Tran, Binh T. Nguyen, Liting Zhou, Cathal Gurrin arxiv

Image-Guided Retrieval with Optional Text (IGROT) unifies visual retrieval (without text) and composed retrieval (with text). Despite its relevance in applications like Google Image and Bing, progress has been limited by the lack of an accessible benchmark and methods that balance performance across subtasks. Large-scale datasets such as MagicLens are comprehensive but computationally prohibitive, while existing models often favor either visual or compositional queries. We introduce FIGROTD, a lightweight yet high-quality IGROT dataset with 16,474 training triplets and 1,262 test triplets across CIR, SBIR, and CSTBIR. To reduce redundancy, we propose the Variance Guided Feature Mask (VaGFeM), which selectively enhances discriminative dimensions based on variance statistics. We further adopt a dual-loss design (InfoNCE + Triplet) to improve compositional reasoning. Trained on FIGROTD, VaGFeM achieves competitive results on nine benchmarks, reaching 34.8 mAP@10 on CIRCO and 75.7 mAP@200 on Sketchy, outperforming stronger baselines despite fewer triplets.

📄 PDF Abstract BibTeX arXiv:2511.22247

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Language-Guided Global Image Editing via Cross-Modal Cyclic Mechanism

2021-01-01 · ICCV 2021 10 · Wentao Jiang, Ning Xu, Jiayun Wang, Chen Gao 외

Editing an image automatically via a linguistic request can significantly save laborious manual work and is friendly to photography novice. In this paper, we focus on the task of language-guided global image editing.…

URIE: Universal Image Enhancement for Visual Recognition in the Wild

2020-07-17 · Taeyoung Son, Juwon Kang, Namyup Kim, Sunghyun Cho 외

Despite the great advances in visual recognition, it has been witnessed that recognition models trained on clean images of common datasets are not robust against distorted images in the real world. To tackle this issue, …

Image Enhancement

URIE: Universal Image Enhancement for Visual Recognition in the Wild

2020-08-01 · ECCV 2020 8 · Taeyoung Son Juwon Kang Namyup Kim Sunghyun Cho Suha Kwak

Despite the great advances in visual recognition, it has been witnessed that recognition models trained on clean images of common datasets are not robust against distorted images in the real world. To tackle this issue, …

Image Enhancement

UniMedSeg: Unified In-Context Learning for Multi-Paradigm 2D/3D Medical Image Segmentation

2026-07-14 · Yunzhou Li, Jiesi Hu, Yanwu Yang, Hanyang Peng 외 arxiv

Medical image segmentation foundation models are expected to generalize across diverse clinical scenarios, yet existing universal methods remain fragmented by prompt paradigms and spatial dimensions. Visual in-context le…

Medical Image SegmentationInteractive Segmentation

MAD: Makeup All-in-One with Cross-Domain Diffusion Model

2025-04-03 · Bo-Kai Ruan, Hong-Han Shuai

Existing makeup techniques often require designing multiple models to handle different inputs and align features across domains for different makeup tasks, e.g., beauty filter, makeup transfer, and makeup removal, leadin…

AllDecoder