paper-with-me

홈 › Papers

EarthBridge: A Solution for 4th Multi-modal Aerial View Image Challenge Translation Track

2026-03-06 · Zhenyuan Chen, Guanyuan Shen, Feng Zhang arxiv

Cross-modal image-to-image translation among Electro-Optical (EO), Infrared (IR), and Synthetic Aperture Radar (SAR) sensors is essential for comprehensive multi-modal aerial-view analysis. However, translating between these modalities is notoriously difficult due to their distinct electromagnetic signatures and geometric characteristics. This paper presents \textbf{EarthBridge}, a high-fidelity translation framework developed for the 4th Multi-modal Aerial View Image Challenge -- Translation (MAVIC-T). We explore two distinct methodologies: \textbf{Diffusion Bridge Implicit Models (DBIM)}, which we generalize using non-Markovian bridge processes for high-quality deterministic sampling, and \textbf{Contrastive Unpaired Translation (CUT)}, which utilizes contrastive learning for structural consistency. Our EarthBridge framework employs a channel-concatenated UNet denoiser trained with Karras-weighted bridge scalings and a specialized "booting noise" initialization to handle the inherent ambiguity in cross-modal mappings. We evaluate these methods across all four challenge tasks (SAR$\rightarrow$EO, SAR$\rightarrow$RGB, SAR$\rightarrow$IR, RGB$\rightarrow$IR), achieving superior spatial detail and spectral accuracy. Our solution achieved a composite score of 0.38, securing the second position on the MAVIC-T leaderboard. Code is available at https://github.com/Bili-Sakura/EarthBridge-Preview.

📄 PDF Abstract BibTeX arXiv:2603.06753

Code (0)

등록된 구현이 없습니다.

Tasks

Image-to-Image TranslationContrastive Learning

Similar Papers 제목 키워드 기반

Cross-View Policy Learning for Street Navigation

2019-06-13 · ICCV 2019 10 · Ang Li, Huiyi Hu, Piotr Mirowski, Mehrdad Farajtabar

The ability to navigate from visual observations in unfamiliar environments is a core component of intelligent agents and an ongoing challenge for Deep Reinforcement Learning (RL). Street View can be a sensible testbed f…

Deep Reinforcement LearningNavigateReinforcement LearningReinforcement Learning (RL)+1

Multiview Aerial Visual Recognition (MAVREC): Can Multi-view Improve Aerial Visual Perception?

2023-12-07 · CVPR 2024 1 · Aritra Dutta, Srijan Das, Jacob Nielsen, Rajatsubhra Chakraborty 외

Despite the commercial abundance of UAVs, aerial data acquisition remains challenging, and the existing Asia and North America-centric open-source UAV datasets are small-scale or low-resolution and lack diversity in scen…

BenchmarkingDiversityobject-detectionObject Detection+1

Multi-Modal Domain Fusion for Multi-modal Aerial View Object Classification

2022-12-14 · Sumanth Udupa, Aniruddh Sikdar, Suresh Sundaram

Object detection and classification using aerial images is a challenging task as the information regarding targets are not abundant. Synthetic Aperture Radar(SAR) images can be used for Automatic Target Recognition(ATR) …

object-detectionObject Detection

Multi-Modal Aerial-Ground Cross-View Place Recognition with Neural ODEs

2025-01-01 · CVPR 2025 1 · Sijie Wang, Rui She, Qiyu Kang, Siqi Li 외

Place recognition (PR) aims at retrieving the query place from a database and plays a crucial role in various applications, including navigation, autonomous driving, and augmented reality. While previous multi-modal …

Autonomous Driving

PBVS 2024 Solution: Self-Supervised Learning and Sampling Strategies for SAR Classification in Extreme Long-Tail Distribution

2024-12-17 · YuHyun Kim, Minwoo Kim, Hyobin Park, Jinwook Jung 외

The Multimodal Learning Workshop (PBVS 2024) aims to improve the performance of automatic target recognition (ATR) systems by leveraging both Synthetic Aperture Radar (SAR) data, which is difficult to interpret but remai…

Self-Supervised Learning