paper-with-me

홈 › Papers

LKFormer: Large Kernel Transformer for Infrared Image Super-Resolution

2024-01-22 · Feiwei Qin, Kang Yan, Changmiao Wang, Ruiquan Ge, Yong Peng, Kai Zhang

Given the broad application of infrared technology across diverse fields, there is an increasing emphasis on investigating super-resolution techniques for infrared images within the realm of deep learning. Despite the impressive results of current Transformer-based methods in image super-resolution tasks, their reliance on the self-attentive mechanism intrinsic to the Transformer architecture results in images being treated as one-dimensional sequences, thereby neglecting their inherent two-dimensional structure. Moreover, infrared images exhibit a uniform pixel distribution and a limited gradient range, posing challenges for the model to capture effective feature information. Consequently, we suggest a potent Transformer model, termed Large Kernel Transformer (LKFormer), to address this issue. Specifically, we have designed a Large Kernel Residual Attention (LKRA) module with linear complexity. This mainly employs depth-wise convolution with large kernels to execute non-local feature modeling, thereby substituting the standard self-attentive layer. Additionally, we have devised a novel feed-forward network structure called Gated-Pixel Feed-Forward Network (GPFN) to augment the LKFormer's capacity to manage the information flow within the network. Comprehensive experimental results reveal that our method surpasses the most advanced techniques available, using fewer parameters and yielding considerably superior performance.The source code will be available at https://github.com/sad192/large-kernel-Transformer.

📄 PDF Abstract BibTeX arXiv:2401.11859

Code (1)

sad192/large-kernel-transformer 공식 구현 pytorch

Tasks

Image Super-ResolutionInfrared image super-resolutionSuper-Resolution

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Residual Connection 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.

Similar Papers 제목 키워드 기반

Multispectral Detection Transformer with Infrared-Centric Sensor Fusion

2025-05-21 · Seongmin Hwang, Daeyoung Han, Moongu Jeon

Multispectral object detection aims to leverage complementary information from visible (RGB) and infrared (IR) modalities to enable robust performance under diverse environmental conditions. In this letter, we propose IC…

Multispectral Object DetectionObjectobject-detectionObject Detection+2

Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics

2026-07-12 · Byung Gyu Chae arxiv

Large language models exhibit remarkable emergent behaviors, yet the physical mechanism governing their collective dynamics remains poorly understood. Cognitive Field Theory predicts that learning reorganizes the collect…

Learning Dynamic Local Context Representations for Infrared Small Target Detection

2024-12-23 · Guoyi Zhang, Guangsheng Xu, Han Wang, Siyang Chen 외

Infrared small target detection (ISTD) is challenging due to complex backgrounds, low signal-to-clutter ratios, and varying target sizes and shapes. Effective detection relies on capturing local contextual information at…

SwinFuse: A Residual Swin Transformer Fusion Network for Infrared and Visible Images

2022-04-25 · Zhishe Wang, Yanlin Chen, Wenyu Shao, Hui Li 외

The existing deep learning fusion methods mainly concentrate on the convolutional neural networks, and few attempts are made with transformer. Meanwhile, the convolutional operation is a content-independent interaction b…

Computational Efficiency

Infrared Small-Dim Target Detection with Transformer under Complex Backgrounds

2021-09-29 · Fangcen Liu, Chenqiang Gao, Fang Chen, Deyu Meng 외

The infrared small-dim target detection is one of the key techniques in the infrared search and tracking system. Since the local regions similar to infrared small-dim targets spread over the whole background, exploring t…