paper-with-me

Papers

Win-Win: Training High-Resolution Vision Transformers from Two Windows

2023-10-01 · Vincent Leroy, Jerome Revaud, Thomas Lucas, Philippe Weinzaepfel

Transformers have become the standard in state-of-the-art vision architectures, achieving impressive performance on both image-level and dense pixelwise tasks. However, training vision transformers for high-resolution pixelwise tasks has a prohibitive cost. Typical solutions boil down to hierarchical architectures, fast and approximate attention, or training on low-resolution crops. This latter solution does not constrain architectural choices, but it leads to a clear performance drop when testing at resolutions significantly higher than that used for training, thus requiring ad-hoc and slow post-processing schemes. In this paper, we propose a novel strategy for efficient training and inference of high-resolution vision transformers. The key principle is to mask out most of the high-resolution inputs during training, keeping only N random windows. This allows the model to learn local interactions between tokens inside each window, and global interactions between tokens from different windows. As a result, the model can directly process the high-resolution input at test time without any special trick. We show that this strategy is effective when using relative positional embedding such as rotary embeddings. It is 4 times faster to train than a full-resolution network, and it is straightforward to use at test time compared to existing approaches. We apply this strategy to three dense prediction tasks with high-resolution data. First, we show on the task of semantic segmentation that a simple setting with 2 windows performs best, hence the name of our method: Win-Win. Second, we confirm this result on the task of monocular depth prediction. Third, we further extend it to the binocular task of optical flow, reaching state-of-the-art performance on the Spring benchmark that contains Full-HD images with an order of magnitude faster inference than the best competitor.

📄 PDF Abstract BibTeX arXiv:2310.00632

Code (0)

등록된 구현이 없습니다.

Tasks

Depth EstimationDepth PredictionOptical Flow EstimationSemantic Segmentation

Methods 이 논문이 사용한 방법론

High-resolution input 설명 없음

Similar Papers 제목 키워드 기반

RSIR Transformer: Hierarchical Vision Transformer using Random Sampling Windows and Important Region Windows

2023-04-13 · Zhemin Zhang, Xun Gong

Recently, Transformers have shown promising performance in various vision tasks. However, the high costs of global self-attention remain challenging for Transformers, especially for high-resolution vision tasks. Local se…

HiT-SR: Hierarchical Transformer for Efficient Image Super-Resolution

2024-07-08 · Xiang Zhang, Yulun Zhang, Fisher Yu

Transformers have exhibited promising performance in computer vision tasks including image super-resolution (SR). However, popular transformer-based SR methods often employ window self-attention with quadratic computatio…

Image Super-ResolutionSuper-Resolution

Vision Big Bird: Random Sparsification for Full Attention

2023-11-10 · Zhemin Zhang, Xun Gong

Recently, Transformers have shown promising performance in various vision tasks. However, the high costs of global self-attention remain challenging for Transformers, especially for high-resolution vision tasks. Inspired…

StyleSwin: Transformer-based GAN for High-resolution Image Generation

2021-12-20 · CVPR 2022 1 · BoWen Zhang, Shuyang Gu, Bo Zhang, Jianmin Bao 외

Despite the tantalizing success in a broad of vision tasks, transformers have not yet demonstrated on-par ability as ConvNets in high-resolution image generative modeling. In this paper, we seek to explore using pure tra…

BlockingComputational EfficiencyGenerative Adversarial NetworkImage Generation+1

Local-to-Global Self-Attention in Vision Transformers

2021-07-10 · Jinpeng Li, Yichao Yan, Shengcai Liao, Xiaokang Yang 외

Transformers have demonstrated great potential in computer vision tasks. To avoid dense computations of self-attentions in high-resolution visual data, some recent Transformer models adopt a hierarchical design, where se…

image-classificationImage ClassificationSemantic Segmentation