Scalable Nested Optimization for Deep Learning

1 · Research Nvidia · Aug. 12, 2026, 7:21 a.m.
Summary
Scalable Nested Optimization for Deep Learning Gradient-based optimization has been critical to the success of machine learning, updating a single set of parameters to minimize a single loss. A growing number of applications rely on a generalization of this, where we have a bilevel or nested optimization of which subsets of parameters update on different objectives nested inside each other. We focus on motivating examples of hyperparameter optimization and generative adversarial networks. Howeve...