Papers
arxiv:2602.16340

The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks

Published on May 24
Authors:
,

Abstract

We study the implicit bias of momentum-based optimizers on smooth homogeneous models. We show that momentum steepest descent algorithms like Muon (spectral norm), MomentumGD (ell_2 norm), and Signum (ell_infty norm) are approximate steepest descent trajectories under a decaying learning rate schedule, proving that these algorithms have a bias towards KKT points of the corresponding margin maximization problem. We extend the analysis to Adam (without the stability constant), which maximizes the ell_infty margin, and to Muon-Signum and Muon-Adam, which maximize a hybrid norm. Our experiments corroborate the theory and show that the identity of the margin maximized depends on the choice of optimizer. Overall, our results extend earlier lines of work on steepest descent in homogeneous models and momentum-based optimizers in linear models.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2602.16340
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2602.16340 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2602.16340 in a dataset README.md to link it from this page.

Spaces citing this paper 1

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.