paper-with-me

홈 › Papers

CFIR: Fast and Effective Long-Text To Image Retrieval for Large Corpora

2024-02-23 · Zijun Long, Xuri Ge, Richard McCreadie, Joemon Jose

Text-to-image retrieval aims to find the relevant images based on a text query, which is important in various use-cases, such as digital libraries, e-commerce, and multimedia databases. Although Multimodal Large Language Models (MLLMs) demonstrate state-of-the-art performance, they exhibit limitations in handling large-scale, diverse, and ambiguous real-world needs of retrieval, due to the computation cost and the injective embeddings they produce. This paper presents a two-stage Coarse-to-Fine Index-shared Retrieval (CFIR) framework, designed for fast and effective large-scale long-text to image retrieval. The first stage, Entity-based Ranking (ER), adapts to long-text query ambiguity by employing a multiple-queries-to-multiple-targets paradigm, facilitating candidate filtering for the next stage. The second stage, Summary-based Re-ranking (SR), refines these rankings using summarized queries. We also propose a specialized Decoupling-BEiT-3 encoder, optimized for handling ambiguous user needs and both stages, which also enhances computational efficiency through vector-based similarity inference. Evaluation on the AToMiC dataset reveals that CFIR surpasses existing MLLMs by up to 11.06% in Recall@1000, while reducing training and retrieval times by 68.75% and 99.79%, respectively. We will release our code to facilitate future research at https://github.com/longkukuhi/CFIR.

📄 PDF Abstract BibTeX arXiv:2402.15276

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyImage RetrievalRe-RankingRetrieval

Similar Papers 제목 키워드 기반

Improving Compactness and Reducing Ambiguity of CFIRE Rule-Based Explanations

2026-01-07 · Sebastian Müller, Tobias Schneider, Ruben Kemna, Vanessa Toborek arxiv

Models trained on tabular data are widely used in sensitive domains, increasing the demand for explanation methods to meet transparency needs. CFIRE is a recent algorithm in this domain that constructs compact surrogate …

SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch

2026-07-29 · Yihao Chen, Shi Chang, Feng Lin, Khaled Chawa 외 hf

LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. Recent benchmarks such as ProgramBench quantify this…

Program Synthesis

A Cross-Font Image Retrieval Network for Recognizing Undeciphered Oracle Bone Inscriptions

2024-09-10 · Zhicong Wu, Qifeng Su, Ke Gu, Xiaodong Shi

Oracle Bone Inscription (OBI) is the earliest mature writing system in China, which represents a crucial stage in the development of hieroglyphs. Nevertheless, the substantial quantity of undeciphered OBI characters rema…

Image RetrievalRetrieval

FAST-GOAL: Fast and Efficient Global-local Object Alignment Learning

2026-05-26 · Hyungyu Choi, Young Kyun Jang, Chanho Eom arxiv

Vision-language models such as CLIP have shown impressive capabilities in aligning images and text, but they often struggle with lengthy and detailed text descriptions due to pre-training on short and concise captions. W…

Computational EfficiencyObject Detection

Atlas: Multi-Scale Attention Improves Long Context Image Modeling

2025-03-16 · Kumar Krishna Agrawal, Long Lian, Longchao Liu, Natalia Harguindeguy 외

Efficiently modeling massive images is a long-standing challenge in machine learning. To this end, we introduce Multi-Scale Attention (MSA). MSA relies on two key ideas, (i) multi-scale representations (ii) bi-directiona…