paper-with-me

홈 › Papers

Bridging Modalities and Tasks: A Unified Hierarchical ViT for SAR-to-Optical Translation and Semantic Segmentation

2026-09-04 · Siyuan Liu, Xuze Zhang, Yongshun Wang, Licong Pan, Hang Liu, Huihui Li arxiv

Synthetic Aperture Radar (SAR) images have all-weather, day-and-night observation capabilities. However, compared with optical images, their speckle noise and non-intuitive scattering mechanism limit the interpretability of the images. Generative models for SAR-to-optical (S2O) conversion can improve visual interpretability, but existing methods often ignore the constraints on semantic structure, which are necessary for downstream tasks, for the sake of visual effects. We propose a unified collaborative dual-task learning framework, termed BMT (Bridging Modalities and Tasks), that jointly optimizes S2O image translation and semantic segmentation through a shared hierarchical Vision Transformer. The framework integrates: (1) a LocalViTBlock that fuses global self-attention with spatial depthwise convolution through a learnable gating mechanism; (2) an enhanced output module combining multi-scale refinement processing, color correction and anti-aliasing, which calibrates channel-level color statistics through feature fusion; (3) a ControlNet-style conditional injection mechanism that encodes SAR wavelet features and segmentation labels into a multi-scale feature pyramid and injects them at each encoder layer through zero-initialized convolution; (4) a bounded Kendall uncertainty weighting scheme that prevents either task from dominating the shared representation. We evaluate the framework under both paired and unpaired translation settings, on the public WHU-OPT-SAR paired dataset and a self-constructed unpaired ship dataset built from HRSID and DIOR, respectively. The experimental results show that the proposed method achieves competitive S2O translation quality and semantic segmentation performance. The dataset and source code have been publicly released at https://github.com/Lewisyuaner/BMT-S2O-main.

📄 PDF Abstract BibTeX arXiv:2609.04726

Code (3)

Tavish9/awesome-daily-AI-arxiv ★ 115
ZhuYingJessica/cv-daily ★ 47
arxivsub/arXivSub_daily_arxiv ★ 4

Tasks

Semantic Segmentation

Similar Papers 제목 키워드 기반

Bridging Visual and Wireless Sensing via a Unified Radiation Field for 3D Radio Map Construction

2026-01-27 · Chaozheng Wen, Jingwen Tong, Zehong Lin, Chenghong Bian 외 arxiv

The emerging applications of next-generation wireless networks demand high-fidelity environmental intelligence. 3D radio maps bridge physical environments and electromagnetic propagation for spectrum planning and environ…

Inverse Rendering

Learning Pixel Trajectories with Multiscale Contrastive Random Walks

2022-01-20 · CVPR 2022 1 · Zhangxing Bian, Allan Jabri, Alexei A. Efros, Andrew Owens

A range of video modeling tasks, from optical flow to multiple object tracking, share the same fundamental challenge: establishing space-time correspondence. Yet, approaches that dominate each space differ. We take a ste…

Multiple Object TrackingObjectObject TrackingOptical Flow Estimation+4

M$^2$CD: A Unified MultiModal Framework for Optical-SAR Change Detection with Mixture of Experts and Self-Distillation

2025-03-25 · Ziyuan Liu, Jiawei Zhang, Wenyu Wang, Yuantao Gu

Most existing change detection (CD) methods focus on optical images captured at different times, and deep learning (DL) has achieved remarkable success in this domain. However, in extreme scenarios such as disaster respo…

Change DetectionDisaster ResponseMixture-of-Experts

High-speed multiwavelength photonic temporal integration using silicon photonics

2025-05-07 · Yi Zhang, Nikolaos Farmakidis, Ioannis Roumpos, Miltiadis Moralis-Pegios 외

Optical systems have been pivotal for energy-efficient computing, performing high-speed, parallel operations in low-loss carriers. While these predominantly analog optical accelerators bypass digitization to perform para…

DSDL: Data Set Description Language for Bridging Modalities and Tasks in AI Data

2024-05-28 · Bin Wang, Linke Ouyang, Fan Wu, Wenchang Ning 외

In the era of artificial intelligence, the diversity of data modalities and annotation formats often renders data unusable directly, requiring understanding and format conversion before it can be used by researchers or d…

Diversity