Zhao Song, Xiaoyu Li, Yang Cao
We lifted 1 functions out of this paper's own repositories and ran 1 of them in a sandbox. "Ran" means the function executed on a synthesized input and returned a value. It is not a reproduction of the paper's results.
| Repository | Role | Ran |
|---|---|---|
| Gunale0926/Grams | — | 1 of 1 |
| Function | Status | Where it lives |
|---|---|---|
| Grams | Ran | Gunale0926/Grams/grams/src/grams.py code served (permissive licence) · get_code("11f3c2d0fa0690aa") |
Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.
We introduce Gradient Descent with Adaptive Momentum Scaling (Grams), a novel optimization algorithm that decouples the direction and magnitude of parameter updates in deep learning. Unlike traditional optimizers that directly integrate momentum into updates, Grams separates the update direction, derived from current gradients, from momentum, which is used solely for adaptive magnitude scaling. This approach enables Grams to achieve improved loss descent compared to state-of-the-art cautious and momentum-based optimizers. We theoretically demonstrate that Grams descents faster than other state-of-the-art optimizers and establish a global convergence guarantee for Grams. We also validate its effectiveness through extensive empirical evaluations. The results demonstrate Grams' superior performance, including faster convergence and better generalization, compared to widely-used optimizers such as Adam, Lion, and their cautious variants. Our results highlight Grams' potential as a transformative approach for efficiently training large language models. Code is available at https://github.com/Gunale0926/Grams.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2412.17107")
get_code_for_paper("2412.17107")
have("2412.17107")
Connect an agent — have() is free.