Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

STDE++: Polynomial-Time Amortization for Linear Differential Operators

Дата публикации: 17-08-2026 20:26:00


Optimizing neural networks with losses that contain high-dimensional and high-order differential operators is expensive to evaluate with backpropagation due to $\mathcal{O}(d^{k})$ scaling of the derivative tensor size and the $\mathcal{O}(2^{k-1}L)$ scaling in the computation graph, where $d$ is the domain dimension, $L$ is the number of ops in the forward computation graph and $k$ is the derivative order. Previous works addressed the polynomial scaling in $d$ by amortizing the computation over the optimization process via randomization. Separately, the exponential scaling in $k$ for univariate functions ($d=1$) was addressed with high-order auto-differentiation (AD). In this work, we show how to efficiently perform arbitrary contractions of the derivative tensor of arbitrary order for multivariate functions by properly constructing the input tangents to univariate high-order AD, which can be used to randomize any differential operator efficiently.
When applied to Physics-Informed Neural Networks (PINNs) and compared against the original PyTorch implementation of SDGD, our method yields about $1.34\times 10^{3}$ average speedup and $31.8\times$ average memory reduction across the three inseparable 100K-dimensional PDEs in our benchmark; the best case is $1.59\times 10^{3}$ speedup and $33.8\times$ memory reduction on Allen-Cahn. We can now solve 1-million-dimensional PDEs in 8 minutes on a single NVIDIA A100 GPU. Furthermore, we proposed new methods for computing mixed partial derivatives using Taylor mode AD, which scales polynomially with the derivative order. This work opens the possibility of using high-order differential operators in large-scale problems.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1 A Natural Primal-Dual Hybrid Gradient Method for Adversarial Neural Network Training on Solving Partial Differential Equation 07.9417-08-2026
2 High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks 08.717-08-2026
3 Statistical guarantees for denoising reflected diffusion models 05.317-08-2026
4 A Common Interface for Automatic Differentiation 09.3217-08-2026
5 Approximation-Free Differentiable Oblique Decision Trees 09.617-08-2026
6 The Sample Complexity of Parameter-Free Stochastic Convex Optimization 05.717-08-2026
7 Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes 07.5217-08-2026
8 Gradient Span Algorithms Make Predictable Progress in High Dimension 06.3817-08-2026
9 End-to-End Deep Learning for Predicting Metric Space-Valued Outputs 010.6617-08-2026
10 Robust training of implicit generative models for multivariate and heavy-tailed distributions with an invariant statistical loss 01017-08-2026

Классификация: Наука. Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 7.12. Источник: jmlr.org.