paper-with-me

홈 › Papers

SARFormer -- An Acquisition Parameter Aware Vision Transformer for Synthetic Aperture Radar Data

2025-04-11 · Jonathan Prexl, Michael Recla, Michael Schmitt

This manuscript introduces SARFormer, a modified Vision Transformer (ViT) architecture designed for processing one or multiple synthetic aperture radar (SAR) images. Given the complex image geometry of SAR data, we propose an acquisition parameter encoding module that significantly guides the learning process, especially in the case of multiple images, leading to improved performance on downstream tasks. We further explore self-supervised pre-training, conduct experiments with limited labeled data, and benchmark our contribution and adaptations thoroughly in ablation experiments against a baseline, where the model is tested on tasks such as height reconstruction and segmentation. Our approach achieves up to 17% improvement in terms of RMSE over baseline models

📄 PDF Abstract BibTeX arXiv:2504.08441

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Position-Wise Feed-Forward Layer 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Multi-Head Attention 설명 없음

Similar Papers 제목 키워드 기반

Local Window Attention Transformer for Polarimetric SAR Image Classification

2023-01-23 · IEEE Geoscience and Remote Sensing Letters 2023 1 · Ali Jamali, Swalpa Kumar Roy, Avik Bhattacharya, Pedram Ghamisi

Convolutional neural networks (CNNs) have recently found great attention in image classification since deep CNNs have exhibited excellent performance in computer vision. Owing to their immense success, of late, scientist…

ClassificationEarth Observationimage-classificationImage Classification

MSRA-SR: Image Super-resolution Transformer with Multi-scale Shared Representation Acquisition

2023-01-01 · ICCV 2023 1 · Xiaoqiang Zhou, Huaibo Huang, Ran He, Zilei Wang 외

Multi-scale feature extraction is crucial for many computer vision tasks, but it is rarely explored in Transformer-based image super-resolution (SR) methods. In this paper, we propose an image super-resolution Transf…

Image Super-ResolutionSuper-Resolution

Multi-Contrast MRI Motion Correction via Parameter-Informed Disentanglement and Adaptive Experts

2026-05-29 · Honglin Xiong, Yuxian Tang, Feng Li, Yulin Wang 외 arxiv

Motion artifacts in magnetic resonance imaging (MRI) degrade diagnostic reliability. Existing deep learning methods are typically contrast-specific and fail to generalize across diverse modalities and artifact severities…

Zero-shot Generalization

ViTOC: Vision Transformer and Object-aware Captioner

2024-11-09 · Feiyang Huang

This paper presents ViTOC (Vision Transformer and Object-aware Captioner), a novel vision-language model for image captioning that addresses the challenges of accuracy and diversity in generated descriptions. Unlike conv…

DiversityImage CaptioningLanguage ModelingLanguage Modelling+1

LUM-ViT: Learnable Under-sampling Mask Vision Transformer for Bandwidth Limited Optical Signal Acquisition

2024-03-03 · Lingfeng Liu, Dong Ni, Hangjie Yuan

Bandwidth constraints during signal acquisition frequently impede real-time detection applications. Hyperspectral data is a notable example, whose vast volume compromises real-time hyperspectral detection. To tackle this…

Binarization