paper-with-me

홈 › Papers

PHA: Patch-Wise High-Frequency Augmentation for Transformer-Based Person Re-Identification

2023-01-01 · CVPR 2023 1 · Guiwei Zhang, Yongfei Zhang, Tianyu Zhang, Bo Li, ShiLiang Pu

Although recent studies empirically show that injecting Convolutional Neural Networks (CNNs) into Vision Transformers (ViTs) can improve the performance of person re-identification, the rationale behind it remains elusive. From a frequency perspective, we reveal that ViTs perform worse than CNNs in preserving key high-frequency components (e.g, clothes texture details) since high-frequency components are inevitably diluted by low-frequency ones due to the intrinsic Self-Attention within ViTs. To remedy such inadequacy of the ViT, we propose a Patch-wise High-frequency Augmentation (PHA) method with two core designs. First, to enhance the feature representation ability of high-frequency components, we split patches with high-frequency components by the Discrete Haar Wavelet Transform, then empower the ViT to take the split patches as auxiliary input. Second, to prevent high-frequency components from being diluted by low-frequency ones when taking the entire sequence as input during network optimization, we propose a novel patch-wise contrastive loss. From the view of gradient optimization, it acts as an implicit augmentation to improve the representation ability of key high-frequency components. This benefits the ViT to capture key high-frequency components to extract discriminative person representations. PHA is necessary during training and can be removed during inference, without bringing extra complexity. Extensive experiments on widely-used ReID datasets validate the effectiveness of our method.

📄 PDF Abstract BibTeX

Code (1)

zhangguiwei610/pha 공식 구현 pytorch

Tasks

Person Re-Identification

Similar Papers 제목 키워드 기반

Full-Frequency Temporal Patching and Structured Masking for Enhanced Audio Classification

2025-08-28 · Aditya Makineni, Baocheng Geng, Qing Tian arxiv

Transformers and State-Space Models (SSMs) have advanced audio classification by modeling spectrograms as sequences of patches. However, existing models such as the Audio Spectrogram Transformer (AST) and Audio Mamba (Au…

Audio Classification

THAT: Token-wise High-frequency Augmentation Transformer for Hyperspectral Pansharpening

2025-08-11 · Hongkun Jin, Hongcheng Jiang, Zejun Zhang, Yuan Zhang 외 arxiv

Transformer-based methods have demonstrated strong potential in hyperspectral pansharpening by modeling long-range dependencies. However, their effectiveness is often limited by redundant token representations and a lack…

Improving Robustness Without Sacrificing Accuracy with Patch Gaussian Augmentation

2019-06-06 · Raphael Gontijo Lopes, Dong Yin, Ben Poole, Justin Gilmer 외

Deploying machine learning systems in the real world requires both high accuracy on clean data and robustness to naturally occurring corruptions. While architectural advances have led to improved accuracy, building robus…

Data Augmentationobject-detectionObject Detection

Data Augmentation in Time Series Forecasting through Inverted Framework

2025-07-15 · Hongming Tan, Ting Chen, Ruochong Jin, Wai Kin Chan

Currently, iTransformer is one of the most popular and effective models for multivariate time series (MTS) forecasting. Thanks to its inverted framework, iTransformer effectively captures multivariate correlation. Howeve…

Data AugmentationTime SeriesTime Series Forecasting

Decoding Human Attentive States from Spatial-temporal EEG Patches Using Transformers

2025-02-06 · Yi Ding, Joon Hei Lee, Shuailei Zhang, Tianze Luo 외

Learning the spatial topology of electroencephalogram (EEG) channels and their temporal dynamics is crucial for decoding attention states. This paper introduces EEG-PatchFormer, a transformer-based deep learning framewor…

Brain Computer InterfaceEEGElectroencephalogram (EEG)