paper-with-me

Papers

REJEPA: A Novel Joint-Embedding Predictive Architecture for Efficient Remote Sensing Image Retrieval

2025-04-04 · Shabnam Choudhury, Yash Salunkhe, Sarthak Mehrotra, Biplab Banerjee

The rapid expansion of remote sensing image archives demands the development of strong and efficient techniques for content-based image retrieval (RS-CBIR). This paper presents REJEPA (Retrieval with Joint-Embedding Predictive Architecture), an innovative self-supervised framework designed for unimodal RS-CBIR. REJEPA utilises spatially distributed context token encoding to forecast abstract representations of target tokens, effectively capturing high-level semantic features and eliminating unnecessary pixel-level details. In contrast to generative methods that focus on pixel reconstruction or contrastive techniques that depend on negative pairs, REJEPA functions within feature space, achieving a reduction in computational complexity of 40-60% when compared to pixel-reconstruction baselines like Masked Autoencoders (MAE). To guarantee strong and varied representations, REJEPA incorporates Variance-Invariance-Covariance Regularisation (VICReg), which prevents encoder collapse by promoting feature diversity and reducing redundancy. The method demonstrates an estimated enhancement in retrieval accuracy of 5.1% on BEN-14K (S1), 7.4% on BEN-14K (S2), 6.0% on FMoW-RGB, and 10.1% on FMoW-Sentinel compared to prominent SSL techniques, including CSMAE-SESD, Mask-VLM, SatMAE, ScaleMAE, and SatMAE++, on extensive RS benchmarks BEN-14K (multispectral and SAR data), FMoW-RGB and FMoW-Sentinel. Through effective generalisation across sensor modalities, REJEPA establishes itself as a sensor-agnostic benchmark for efficient, scalable, and precise RS-CBIR, addressing challenges like varying resolutions, high object density, and complex backgrounds with computational efficiency.

📄 PDF Abstract BibTeX arXiv:2504.03169

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyContent-Based Image RetrievalImage RetrievalRetrieval

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

HQ-JEPA: Hybrid Quantum Joint-Embedding Predictive Architecture for Cross-Modal Remote Sensing Representation Learning

2026-05-29 · Md Aminur Hossain, Ayush V. Patel, Sanjay K. Singh, Biplab Banerjee arxiv

We introduce HQ-JEPA, a hybrid quantum-classical joint-embedding predictive architecture for cross-modal remote sensing representation learning. The proposed framework extends JEPA-style masked latent prediction to paire…

Representation Learning

CR-JEPA: Cross-Modal Joint-Embedding Predictive Learning for Remote Sensing Image Retrieval

2026-05-30 · Md Aminur Hossain, Ayush V. Patel, Nitant Dube, Biplab Banerjee arxiv

Cross-modal remote sensing image retrieval aims to retrieve semantically related scenes across heterogeneous sensing modalities. This remains challenging because paired observations may differ substantially in imaging ph…

Cross-Modal RetrievalImage Retrieval

Hierarchical JEPA Meets Predictive Remote Control in Beyond 5G Networks

2026-01-28 · Abanoub M. Girgis, Ibtissam Labriji, Mehdi Bennis arxiv

In wireless networked control systems, ensuring timely and reliable state updates from distributed devices to remote controllers is essential for robust control performance. However, when multiple devices transmit high-d…

Coupled Control and Wireless World Models for Resilient Remote Robotic Control

2026-09-04 · H. P. Madushanka, Sumudu Samarakoon, Mehdi Bennis arxiv

Remote robotic systems operating over wireless networks must maintain reliable control despite limited communication resources, changing channel conditions, and environmental disturbances.However, continuously transmitti…

Time-Series JEPA for Predictive Remote Control under Capacity-Limited Networks

2024-06-07 · Abanoub M. Girgis, Alvaro Valcarce, Mehdi Bennis

In remote control systems, transmitting large data volumes (e.g. video feeds) from wireless sensors to faraway controllers is challenging when the uplink channel capacity is limited (e.g. RedCap devices or massive wirele…

Self-Supervised LearningTime Series