The Rich and the Simple: On the Implicit Bias of Adam and SGD
bayes' optimal predictorbinary classificationdecision boundarydistribution shiftsfirst-order methodsgaussian datageneralizationimplicit biasoptimization algorithmpopulation gradientssimplicity biasspurious correlationsstochastic gradient descenttest accuracy
Adam is the de facto optimization algorithm for several deep learning applications, but an understanding of its implicit bias and how it differs from other algorithms, particularly standard first-order methods such as (stochastic) gradient descent (GD), remains limited. In practice, neural networks (NNs) trained with SGD are known to exhibit simplicity bias