Fourier Image Transformer
Transformer architectures show spectacular performance on NLP tasks and have recently also been used for tasks such as image completion or image classification. Here we propose to use a sequential image representation, where each prefix of the complete sequence describes the whole image at reduced resolution. Using such Fourier Domain Encodings (FDEs), an auto-regressive image completion task is equivalent to predicting a higher resolution output given a low-resolution input. Additionally, we show that an encoder-decoder setup can be used to query arbitrary Fourier coefficients given a set of Fourier domain observations. We demonstrate the practicality of this approach in the context of computed tomography (CT) image reconstruction. In summary, we show that Fourier Image Transformer (FIT) can be used to solve relevant image analysis tasks in Fourier space, a domain inherently inaccessible to convolutional architectures.
Code (1)
Tasks
Computed Tomography (CT)Decoderimage-classificationImage ClassificationImage ReconstructionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
FourierSR: A Fourier Token-based Plugin for Efficient Image Super-Resolution
Image super-resolution (SR) aims to recover low-resolution images to high-resolution images, where improving SR efficiency is a high-profile challenge. However, commonly used units in SR, like convolutions and window-bas…
Image Super-ResolutionSuper-ResolutionTransformer with Fourier Integral Attentions
Multi-head attention empowers the recent success of transformers, the state-of-the-art models that have achieved remarkable success in sequence modeling and beyond. These attention mechanisms compute the pairwise dot pro…
image-classificationImage ClassificationLanguage ModelingLanguage Modelling+1Bootstrapping Interactive Image-Text Alignment for Remote Sensing Image Captioning
Recently, remote sensing image captioning has gained significant attention in the remote sensing community. Due to the significant differences in spatial resolution of remote sensing images, existing methods in this fiel…
Causal Language ModelingContrastive LearningImage CaptioningLanguage Modeling+3F2former: When Fractional Fourier Meets Deep Wiener Deconvolution and Selective Frequency Transformer for Image Deblurring
Recent progress in image deblurring techniques focuses mainly on operating in both frequency and spatial domains using the Fourier transform (FT) properties. However, their performance is limited due to the dependency of…
DeblurringDecoderImage DeblurringImage RestorationContextual Learning in Fourier Complex Field for VHR Remote Sensing Images
Very high-resolution (VHR) remote sensing (RS) image classification is the fundamental task for RS image analysis and understanding. Recently, transformer-based models demonstrated outstanding potential for learning high…
Classificationimage-classificationImage Classification