paper-with-me

홈 › Papers

Learning Faster without Deeper Networks: A*-Inspired Batch Selection for Efficient CNN Training

2026-07-17 · Anxhelo Shehu, Enes Stastoli, Arben Cela arxiv

Common practice when training Convolutional Neural Networks (CNNs) is to use randomly shuffled mini-batches. This creates two limitations: slower convergence, and a diminishing learning signal, since many samples are quickly classified as easy during training. We address these inefficiencies with A*-Inspired Batch Selection (A*-BS), a lightweight, model-agnostic strategy that formulates mini-batch scheduling as a heuristic search problem. Each batch is treated as a node in a search space and ranked using an A*-like score combining a loss-based difficulty measure with a reuse penalty. This encourages informative gradient updates and batch diversity throughout training, without modifying network architectures or optimization algorithms, so it integrates seamlessly into existing pipelines. We evaluate A*-BS on the twelve 2D classification tasks of the MedMNIST-v2 benchmark, using a deliberately simple architecture of approximately 2.25x10^5 parameters, compared against the ResNet-18 and ResNet-50 baselines reported by the benchmark. On half of these tasks, the lightweight model with A*-BS reaches higher accuracy and AUC than both ResNet baselines, with relative gains of up to 15%. An ablation under identical architecture and hyperparameters shows A*-BS outperforms random batch shuffling on all twelve tasks. Wall-clock measurements further show the lightweight CNN with A*-BS trains substantially faster than ResNet-18 and ResNet-50 on identical hardware. These results indicate that intelligent batch ordering can partially compensate for reduced architectural complexity, offering a computationally efficient alternative to deeper models, with reliability reinforced by strong performance even against deeper, more sophisticated architectures.

📄 PDF Abstract BibTeX arXiv:2607.15745

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Comparing Normalization Methods for Limited Batch Size Segmentation Neural Networks

2020-11-23 · Martin Kolarik, Radim Burget, Kamil Riha

The widespread use of Batch Normalization has enabled training deeper neural networks with more stable and faster results. However, the Batch Normalization works best using large batch size during training and as the sta…

Segmentation

The Effect of Network Width on the Performance of Large-batch Training

2018-06-11 · NeurIPS 2018 12 · Lingjiao Chen, Hongyi Wang, Jinman Zhao, Dimitris Papailiopoulos 외

Distributed implementations of mini-batch stochastic gradient descent (SGD) suffer from communication overheads, attributed to the high frequency of gradient updates inherent in small-batch training. Training with large …

Grappa: Gradient-Only Communication for Scalable Graph Neural Network Training

2026-02-02 · Chongyang Xu, Christoph Siebenbrunner, Laurent Bindschaedler arxiv

Cross-partition edges dominate the cost of distributed GNN training: fetching remote features and activations per iteration overwhelms the network as graphs deepen and partition counts grow. Grappa is a distributed GNN t…

Graph Neural Network

Online Batch Selection for Faster Training of Neural Networks

2015-11-19 · Ilya Loshchilov, Frank Hutter

Deep neural networks are commonly trained using stochastic non-convex optimization procedures, which are driven by gradient information estimated on fractions (batches) of the dataset. While it is commonly accepted that …

Multi-Label Adaptive Batch Selection by Highlighting Hard and Imbalanced Samples

2024-03-27 · Ao Zhou, Bin Liu, Jin Wang, Grigorios Tsoumakas

Deep neural network models have demonstrated their effectiveness in classifying multi-label data from various domains. Typically, they employ a training mode that combines mini-batches with optimizers, where each sample …