Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

Optimization and Generalization of Gradient Descent for Shallow ReLU Networks with Minimal Width

Дата публикации: 17-08-2026 20:26:00


Understanding the generalization and optimization of neural networks is a longstanding problem in modern learning theory. The prior analysis often leads to risk bounds of order $1/\sqrt{n}$ for ReLU networks, where $n$ is the sample size. In this paper, we present a general optimization and generalization analysis for gradient descent applied to shallow ReLU networks. We develop convergence rates of the order $1/T$ for gradient descent with $T$ iterations, and show that the gradient descent iterates fall inside local balls around either an initialization point or a reference point. Then we develop improved Rademacher complexity estimates by using the activation pattern of the ReLU function in these local balls. We apply our general result to NTK-separable data with a margin $\gamma$, and develop an almost optimal risk bound of the order $1/(n\gamma^2)$ for the ReLU network with a polylogarithmic width.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1 Beyond Unconstrained Features: Neural Collapse for Shallow Neural Networks with General Data 04.3617-08-2026
2 High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks 08.717-08-2026
3 A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks 010.9717-08-2026
4 Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes 07.5217-08-2026
5 Optimal Approximation and Generalization Errors for Deep Convolutional Neural Networks 07.3117-08-2026
6 Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks 0517-08-2026
7 A Mean-Field Analysis of Neural Stochastic Gradient Descent-Ascent for Functional Minimax Optimization 09.8217-08-2026
8 Stochastic Gradient Methods: Bias, Stability and Generalization 06.317-08-2026
9 Reparameterized Complex-valued Neurons Can Efficiently Learn More than Real-valued Neurons via Gradient Descent 011.8617-08-2026
10 Statistical Learning Theory for Neural Operators 010.2117-08-2026

Классификация: . Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 3.84. Источник: jmlr.org.