paper-with-me

홈 › Papers

Cream of the Crop: Distilling Prioritized Paths For One-Shot Neural Architecture Search

2020-10-29 · NeurIPS 2020 12 · Houwen Peng, Hao Du, Hongyuan Yu, Qi Li, Jing Liao, Jianlong Fu

One-shot weight sharing methods have recently drawn great attention in neural architecture search due to high efficiency and competitive performance. However, weight sharing across models has an inherent deficiency, i.e., insufficient training of subnetworks in hypernetworks. To alleviate this problem, we present a simple yet effective architecture distillation method. The central idea is that subnetworks can learn collaboratively and teach each other throughout the training process, aiming to boost the convergence of individual models. We introduce the concept of prioritized path, which refers to the architecture candidates exhibiting superior performance during training. Distilling knowledge from the prioritized paths is able to boost the training of subnetworks. Since the prioritized paths are changed on the fly depending on their performance and complexity, the final obtained paths are the cream of the crop. We directly select the most promising one from the prioritized paths as the final architecture, without using other complex search methods, such as reinforcement learning or evolution algorithms. The experiments on ImageNet verify such path distillation method can improve the convergence ratio and performance of the hypernetwork, as well as boosting the training of subnetworks. The discovered architectures achieve superior performance compared to the recent MobileNetV3 and EfficientNet families under aligned settings. Moreover, the experiments on object detection and more challenging search space show the generality and robustness of the proposed method. Code and models are available at https://github.com/microsoft/cream.git.

📄 PDF Abstract BibTeX arXiv:2010.15821

Code (2)

microsoft/cream 공식 구현 pytorch
LibrarristShalinward/while_true_involute-Cream_of_the_Crop paddle

Tasks

Neural Architecture Searchobject-detectionObject Detection

Methods 이 논문이 사용한 방법론

Pointwise Convolution Pointwise Convolution is a type of convolution that uses a 1x1 kernel: a kernel that iterates through every single point. This…
Depthwise Convolution Depthwise Convolution is a type of convolution where we apply a single convolutional filter for each input channel. In the regular 2D…
ReLU6 ReLU6 is a modification of the rectified linear unit where we limit the activation to a maximum size of $6$. This is due to increased…
Depthwise Separable Convolution While standard convolution performs the channelwise and spatial-wise computation in one step, Depthwise Separable Convolution …
Batch Normalization 설명 없음
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Inverted Residual Block 설명 없음
Hard Swish Hard Swish is a type of activation function based on Swish, but replaces the computationally expensive sigmoid with a piecewise…

Similar Papers 제목 키워드 기반

Multi-robot Path Planning in Well-formed Infrastructures: Prioritized Planning vs. Prioritized Wait Adjustment (Preliminary Results)

2018-07-05 · Anton Andreychuk, Konstantin Yakovlev

We study the problem of planning collision-free paths for a group of homogeneous robots. We propose a novel approach for turning the paths that were planned egocentrically by the robots, e.g. without taking other robots'…

Unsupervised vs. transfer learning for multimodal one-shot matching of speech and images

2020-08-14 · Leanne Nortje, Herman Kamper

We consider the task of multimodal one-shot speech-image matching. An agent is shown a picture along with a spoken word describing the object in the picture, e.g. cookie, broccoli and ice-cream. After observing one paire…

Transfer Learning

The Cream Rises to the Top: Efficient Reranking Method for Verilog Code Generation

2025-09-24 · Guang Yang, Wei Zheng, Xiang Chen, Yifan Sun 외 arxiv

LLMs face significant challenges in Verilog generation due to limited domain-specific knowledge. While sampling techniques improve pass@k metrics, hardware engineers need one trustworthy solution rather than uncertain ca…

Code Generation

StereoDistill: Pick the Cream from LiDAR for Distilling Stereo-based 3D Object Detection

2023-01-04 · Zhe Liu, Xiaoqing Ye, Xiao Tan, Errui Ding 외

In this paper, we propose a cross-modal distillation method named StereoDistill to narrow the gap between the stereo and LiDAR-based approaches via distilling the stereo detectors from the superior LiDAR model at the res…

3D Object Detectionobject-detectionObject Detection

Efficient View Planning Guided by Previous-Session Reconstruction for Repeated Plant Monitoring

2025-10-08 · Sicong Pan, Luca Lobefaro, Moein Taherkhani, Xuying Huang 외 arxiv

Repeated plant monitoring is essential for tracking crop growth, and 3D reconstruction enables consistent comparison across monitoring sessions. However, rebuilding a 3D model from scratch in every session is costly and …

3D Reconstruction