OrScale Optimizes Layer-Wise Scaling for Orthogonalized Optimizers
September 1, 2026
OrScale introduces a dynamic per-layer scalar for orthogonalized optimization by adapting LARS/LAMB trust-ratio principles. It allows the Moonlight recipe to transfer to Muon-based training without additional hyperparameter sweeps while maintaining an O(1/sqrt(T)) convergence rate.
HOW THIS AFFECTS YOU
●
researcherYou can now adapt orthogonalized optimization methods to different training recipes without exhaustive hyperparameter tuning.