paper-with-me

홈 › Papers

SSBNet: Improving Visual Recognition Efficiency by Adaptive Sampling

2022-07-23 · Ho Man Kwan, Shenghui Song

Downsampling is widely adopted to achieve a good trade-off between accuracy and latency for visual recognition. Unfortunately, the commonly used pooling layers are not learned, and thus cannot preserve important information. As another dimension reduction method, adaptive sampling weights and processes regions that are relevant to the task, and is thus able to better preserve useful information. However, the use of adaptive sampling has been limited to certain layers. In this paper, we show that using adaptive sampling in the building blocks of a deep neural network can improve its efficiency. In particular, we propose SSBNet which is built by inserting sampling layers repeatedly into existing networks like ResNet. Experiment results show that the proposed SSBNet can achieve competitive image classification and object detection performance on ImageNet and COCO datasets. For example, the SSB-ResNet-RS-200 achieved 82.6% accuracy on ImageNet dataset, which is 0.6% higher than the baseline ResNet-RS-152 with a similar complexity. Visualization shows the advantage of SSBNet in allowing different layers to focus on different positions, and ablation studies further validate the advantage of adaptive sampling over uniform methods.

📄 PDF Abstract BibTeX arXiv:2207.11511

Code (0)

등록된 구현이 없습니다.

Tasks

Dimensionality Reductionimage-classificationImage Classificationobject-detectionObject Detection

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Kaiming Initialization 설명 없음
Bottleneck Residual Block A Bottleneck Residual Block is a variant of the residual block that utilises 1x1 convolutions to create a bottleneck. The…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Residual Connection 설명 없음
Average Pooling 설명 없음
Batch Normalization 설명 없음

Similar Papers 제목 키워드 기반

Efficient Human Vision Inspired Action Recognition using Adaptive Spatiotemporal Sampling

2022-07-12 · Khoi-Nguyen C. Mac, Minh N. Do, Minh P. Vo

Adaptive sampling that exploits the spatiotemporal redundancy in videos is critical for always-on action recognition on wearable devices with limited computing and battery resources. The commonly used fixed sampling stra…

Action Recognition

MRS-VPR: a multi-resolution sampling based global visual place recognition method

2019-02-26 · Peng Yin, Rangaprasad Arun Srivatsan, Yin Chen, Xueqian Li 외

Place recognition and loop closure detection are challenging for long-term visual navigation tasks. SeqSLAM is considered to be one of the most successful approaches to achieving long-term localization under varying envi…

Loop Closure DetectionVisual NavigationVisual Place Recognition

3D Adaptive Structural Convolution Network for Domain-Invariant Point Cloud Recognition

2024-07-05 · Younggun Kim, Beomsik Cho, Seonghoon Ryoo, Soomok Lee

Adapting deep learning networks for point cloud data recognition in self-driving vehicles faces challenges due to the variability in datasets and sensor technologies, emphasizing the need for adaptive techniques to maint…

Motion-driven Visual Tempo Learning for Video-based Action Recognition

2022-02-24 · TIP 2022 5 · Yuanzhong Liu, Junsong Yuan, Zhigang Tu

Action visual tempo characterizes the dynamics and the temporal scale of an action, which is helpful to distinguish human actions that share high similarities in visual dynamics and appearance. Previous methods capture t…

Action Recognition

Task-adaptive Spatial-Temporal Video Sampler for Few-shot Action Recognition

2022-07-20 · Huabin Liu, Weixian Lv, John See, Weiyao Lin

A primary challenge faced in few-shot action recognition is inadequate video data for training. To address this issue, current methods in this field mainly focus on devising algorithms at the feature level while little a…

Action RecognitionFew-Shot action recognitionFew Shot Action Recognition