paper-with-me

Papers

Beyond Pixels: A Training-Free, Text-to-Text Framework for Remote Sensing Image Retrieval

2025-12-11 · J. Xiao, Y. Guo, X. Zi, K. Thiyagarajan, C. Moreira, M. Prasad arxiv

Semantic retrieval of remote sensing (RS) images is a critical task fundamentally challenged by the \textquote{semantic gap}, the discrepancy between a model's low-level visual features and high-level human concepts. While large Vision-Language Models (VLMs) offer a promising path to bridge this gap, existing methods often rely on costly, domain-specific training, and there is a lack of benchmarks to evaluate the practical utility of VLM-generated text in a zero-shot retrieval context. To address this research gap, we introduce the Remote Sensing Rich Text (RSRT) dataset, a new benchmark featuring multiple structured captions per image. Based on this dataset, we propose a fully training-free, text-only retrieval reference called TRSLLaVA. Our methodology reformulates cross-modal retrieval as a text-to-text (T2T) matching problem, leveraging rich text descriptions as queries against a database of VLM-generated captions within a unified textual embedding space. This approach completely bypasses model training or fine-tuning. Experiments on the RSITMD and RSICD benchmarks show our training-free method is highly competitive with state-of-the-art supervised models. For instance, on RSITMD, our method achieves a mean Recall of 42.62\%, nearly doubling the 23.86\% of the standard zero-shot CLIP baseline and surpassing several top supervised models. This validates that high-quality semantic representation through structured text provides a powerful and cost-effective paradigm for remote sensing image retrieval.

📄 PDF Abstract BibTeX arXiv:2512.10596

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalSemantic RetrievalImage Retrieval

Results from the Paper

RankTaskDatasetModelMetrics
#31 Cross-Modal Retrieval RSITMD Remote Mean Recall: 42.62

Similar Papers 제목 키워드 기반

World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible

2026-06-11 · Hao Zhang, Mohamed El Banani, Jen-Hao Cheng, Paul Zhang 외 arxiv

Image-to-3D methods often trade off faithfulness and completeness: depth estimators are anchored to input pixels but stop at the visible surface, while image-to-3D models generate complete shapes that are often misaligne…

3D scene Editing

Learning to Decipher from Pixels: A Case Study of Copiale

2026-04-26 · Lei Kang, Giuseppe De Gregorio, Raphaela Heil, Alicia Fornés 외 arxiv

Historical encrypted manuscripts require both paleographic interpretation of cipher symbols and cryptanalytic recovery of plaintext. Most existing computational workflows rely on a transcription-first paradigm, in which …

AMLRIS: Alignment-aware Masked Learning for Referring Image Segmentation

2026-02-26 · Tongfei Chen, Shuo Yang, Yuguang Yang, Linlin Yang 외 arxiv

Referring Image Segmentation (RIS) aims to segment the object in an image uniquely referred to by a natural language expression. However, RIS training often contains hard-to-align and instance-specific visual signals; op…

Image Segmentation

Mining Contextual Information Beyond Image for Semantic Segmentation

2021-08-26 · ICCV 2021 10 · Zhenchao Jin, Tao Gong, Dongdong Yu, Qi Chu 외

This paper studies the context aggregation problem in semantic image segmentation. The existing researches focus on improving the pixel representations by aggregating the contextual information within individual images. …

Image SegmentationSegmentationSemantic Segmentation

Improved Stochastic Texture Filtering Through Sample Reuse

2025-04-07 · Bartlomiej Wronski, Matt Pharr, Tomas Akenine-Möller

Stochastic texture filtering (STF) has re-emerged as a technique that can bring down the cost of texture filtering of advanced texture compression methods, e.g., neural texture compression. However, during texture magnif…

Denoising