paper-with-me

홈 › Papers

Guaranteed Recovery of One-Hidden-Layer Neural Networks via Cross Entropy

2018-02-18 · ICLR 2019 5 · Haoyu Fu, Yuejie Chi, Yingbin Liang

We study model recovery for data classification, where the training labels are generated from a one-hidden-layer neural network with sigmoid activations, also known as a single-layer feedforward network, and the goal is to recover the weights of the neural network. We consider two network models, the fully-connected network (FCN) and the non-overlapping convolutional neural network (CNN). We prove that with Gaussian inputs, the empirical risk based on cross entropy exhibits strong convexity and smoothness {\em uniformly} in a local neighborhood of the ground truth, as soon as the sample complexity is sufficiently large. This implies that if initialized in this neighborhood, gradient descent converges linearly to a critical point that is provably close to the ground truth. Furthermore, we show such an initialization can be obtained via the tensor method. This establishes the global convergence guarantee for empirical risk minimization using cross entropy via gradient descent for learning one-hidden-layer neural networks, at the near-optimal sample and computational complexity with respect to the network input dimension without unrealistic assumptions such as requiring a fresh set of samples at each iteration.

📄 PDF Abstract BibTeX arXiv:1802.06463

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A First-Order Mean Field Control Analysis of Transformer Layers under Cross-Entropy Training

2026-06-22 · Cheng Huan, Hongwei Yuan arxiv

We study Transformer-type residual layers under cross-entropy training through a continuous-depth mean field control viewpoint. Depth is treated as time, layer parameters as controls, and the residual Transformer recursi…

Recovery Guarantees for One-hidden-layer Neural Networks

2017-06-10 · ICML 2017 8 · Kai Zhong, Zhao Song, Prateek Jain, Peter L. Bartlett 외

In this paper, we consider regression problems with one-hidden-layer neural networks (1NNs). We distill some properties of activation functions that lead to $\mathit{local~strong~convexity}$ in the neighborhood of the gr…

Learning One-hidden-layer Neural Networks on Gaussian Mixture Models with Guaranteed Generalizability

2021-01-01 · Hongkang Li, Shuai Zhang, Meng Wang

We analyze the learning problem of fully connected neural networks with the sigmoid activation function for binary classification from the setup of model estimation. The outputs are assumed to be generated by a ground-tr…

Binary Classification

Where Should Diffusion Enter a Language Model? Geometry-Guided Hidden-State Replacement

2026-05-14 · Injin Kong, Hyoungjoon Lee, Yohan Jo arxiv

Continuous diffusion language models lag behind autoregressive transformers, partly because diffusion is applied in spaces poorly suited to language denoising and token recovery. We propose DiHAL, a geometry-guided diffu…

Is There Any Recovery Guarantee with Coupled Structured Matrix Factorization for Hyperspectral Super-Resolution?

2019-07-30

Coupled structured matrix factorization (CoSMF) for hyperspectral super-resolution (HSR) has recently drawn significant interest in hyperspectral imaging for remote sensing. Presently there is very few work that studies …

Super-Resolution