paper-with-me

Papers

Submodular Mini-Batch Training in Generative Moment Matching Networks

2017-07-18 · Jun Qi

This article was withdrawn because (1) it was uploaded without the co-authors' knowledge or consent, and (2) there are allegations of plagiarism.

📄 PDF Abstract BibTeX arXiv:1707.05721

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Submodular Batch Selection for Training Deep Neural Networks

2019-06-20 · K J Joseph, Vamshi Teja R, Krishnakant Singh, Vineeth N. Balasubramanian

Mini-batch gradient descent based methods are the de facto algorithms for training neural network architectures today. We introduce a mini-batch selection strategy based on submodular function maximization. Our novel sub…

Combinatorial OptimizationDiversityInformativeness

Optimal approximation for unconstrained non-submodular minimization

2019-05-29 · ICML 2020 1 · Marwa El Halabi, Stefanie Jegelka

Submodular function minimization is well studied, and existing algorithms solve it exactly or up to arbitrary accuracy. However, in many applications, such as structured sparse learning or batch Bayesian optimization, th…

Bayesian OptimizationSparse Learning

Perfect Parallelization in Mini-Batch SGD with Classical Momentum Acceleration

2026-05-18 · Sachin Garg, Michał Dereziński arxiv

Accelerating stochastic gradient methods with classical momentum schemes, such as Polyak's heavy ball, has proven highly successful in training large-scale machine learning models, particularly when combined with the har…

Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum

2025-01-15 · Keisuke Kamo, Hideaki Iiduka

Stochastic gradient descent with momentum (SGDM), which is defined by adding a momentum term to SGD, has been well studied in both theory and practice. Theoretically investigated results showed that the settings of the l…

Co-Mixup: Saliency Guided Joint Mixup with Supermodular Diversity

2021-02-05 · ICLR 2021 1 · Jang-Hyun Kim, Wonho Choo, Hosan Jeong, Hyun Oh Song

While deep neural networks show great performance on fitting to the training distribution, improving the networks' generalization performance to the test distribution and robustness to the sensitivity to input perturbati…

Diversity