paper-with-me

홈 › Papers

HRFormer: High-Resolution Transformer for Dense Prediction

2021-10-18 · Yuhui Yuan, Rao Fu, Lang Huang, WeiHong Lin, Chao Zhang, Xilin Chen, Jingdong Wang

We present a High-Resolution Transformer (HRFormer) that learns high-resolution representations for dense prediction tasks, in contrast to the original Vision Transformer that produces low-resolution representations and has high memory and computational cost. We take advantage of the multi-resolution parallel design introduced in high-resolution convolutional networks (HRNet), along with local-window self-attention that performs self-attention over small non-overlapping image windows, for improving the memory and computation efficiency. In addition, we introduce a convolution into the FFN to exchange information across the disconnected image windows. We demonstrate the effectiveness of the High-Resolution Transformer on both human pose estimation and semantic segmentation tasks, e.g., HRFormer outperforms Swin transformer by $1.3$ AP on COCO pose estimation with $50\%$ fewer parameters and $30\%$ fewer FLOPs. Code is available at: https://github.com/HRNet/HRFormer.

📄 PDF Abstract BibTeX arXiv:2110.09408

Code (1)

HRNet/HRFormer 공식 구현 pytorch

Tasks

Image ClassificationMulti-Person Pose EstimationPose EstimationPredictionSemantic SegmentationVocal Bursts Intensity Prediction

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Stochastic Depth Stochastic Depth aims to shrink the depth of a network during training, while keeping it unchanged during testing. This is achieved by randomly dropping entire…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

HRFormer: High-Resolution Vision Transformer for Dense Predict

2021-12-01 · NeurIPS 2021 12 · Yuhui Yuan, Rao Fu, Lang Huang, WeiHong Lin 외

We present a High-Resolution Transformer (HRFormer) that learns high-resolution representations for dense prediction tasks, in contrast to the original Vision Transformer that produces low-resolution representations and …

Pose EstimationSemantic SegmentationVocal Bursts Intensity Prediction

Multi-Modal Conditioned High-Resolution Transformer for Urban Electromagnetic Field Map Prediction Download PDF

2026-06-26 · Do-Eon Kim, Dongryul Park, Seungyoung Ahn, Namwoo Kang 외 arxiv

Predicting electromagnetic field (EMF) strength in urban environments is essential for cellular network planning but computationally expensive with physics-based simulators. We propose a multi-conditioned dense predictio…

HRTransNet: HRFormer-Driven Two-Modality Salient Object Detection

2023-01-08 · Bin Tang, Zhengyi Liu, Yacheng Tan, Qian He

The High-Resolution Transformer (HRFormer) can maintain high-resolution representation and share global receptive fields. It is friendly towards salient object detection (SOD) in which the input and output have the same …

global-optimizationObjectobject-detectionObject Detection+2

AggPose: Deep Aggregation Vision Transformer for Infant Pose Estimation

2022-05-11 · Xu Cao, Xiaoye Li, Liya Ma, Yi Huang 외

Movement and pose assessment of newborns lets experienced pediatricians predict neurodevelopmental disorders, allowing early intervention for related diseases. However, most of the newest AI approaches for human pose est…

Keypoint DetectionPose Estimation

DEHRFormer: Real-time Transformer for Depth Estimation and Haze Removal from Varicolored Haze Scenes

2023-03-13 · Sixiang Chen, Tian Ye, Jun Shi, Yun Liu 외

Varicolored haze caused by chromatic casts poses haze removal and depth estimation challenges. Recent learning-based depth estimation methods are mainly targeted at dehazing first and estimating depth subsequently from h…

Contrastive LearningDepth Estimation