The Case for Full-Matrix Adaptive Regularization

Abstract

Adaptive regularization methods come in diagonal and full-matrix variants.However, only the former have enjoyed widespread adoption in traininglarge-scale deep models. This is due to the computational overhead ofmanipulating a full matrix in high dimension. In this paper, we show how tomake full-matrix adaptive regularization practical and useful. We present GGT,a truly scalable full-matrix adaptive optimizer. At the heart of our algorithmis an efficient method for computing the inverse square root of a low-rankmatrix. We show that GGT converges to first-order local minima, providing thefirst rigorous theoretical analysis of adaptive regularization in non-convexoptimization. In preliminary experiments, GGT trains faster across a variety ofsynthetic tasks and standard deep learning benchmarks.

Quick Read (beta)

loading the full paper ...