Optimisers.jl
Installation: OptimizationOptimisers.jl
To use this package, install the OptimizationOptimisers package:
import Pkg;
Pkg.add("OptimizationOptimisers");In addition to the optimisation algorithms provided by the Optimisers.jl package this subpackage also provides the Sophia optimisation algorithm.
Reexported Optimisers.jl API
using OptimizationOptimisers brings Optimisers.jl's rules into scope, so that solve(prob, Adam(0.05)) works without a separate using Optimisers. These names are owned and documented by Optimisers.jl; this package only re-exports them.
- Gradient descent rules:
Descent,Momentum,Nesterov,Rprop - Adaptive rules:
RMSProp,Adam,RAdam,AdaMax,OAdam,AdaGrad,AdaDelta,AMSGrad,NAdam,AdamW,AdaBelief,Lion - Legacy all-caps spellings kept for compatibility:
ADAM,ADAMW,RADAM,OADAM,NADAM,ADAGrad,ADADelta - Gradient modifiers and combinators:
ClipGrad,ClipNorm,WeightDecay,SignDecay,AccumGrad,OptimiserChain - The
Optimisersmodule itself; the rule supertype isOptimisers.AbstractRule
Optimisers' own optimiser-driving interface — Optimisers.setup, Optimisers.update, Optimisers.update!, Optimisers.apply!, Optimisers.destructure, Optimisers.trainables — is deliberately not re-exported: solve drives the rule for you, and Optimisers.init would collide with the SciML init.
Anything else from Optimisers.jl must be imported from Optimisers directly.
List of optimizers
Optimisers.Descent: Classic gradient descent optimizer with learning ratesolve(problem, Descent(η))ηis the learning rateDefaults:
η = 0.1
Optimisers.Momentum: Classic gradient descent optimizer with learning rate and momentumsolve(problem, Momentum(η, ρ))ηis the learning rateρis the momentumDefaults:
η = 0.01ρ = 0.9
Optimisers.Nesterov: Gradient descent optimizer with learning rate and Nesterov momentumsolve(problem, Nesterov(η, ρ))ηis the learning rateρis the Nesterov momentumDefaults:
η = 0.01ρ = 0.9
Optimisers.RMSProp: RMSProp optimizersolve(problem, RMSProp(η, ρ))ηis the learning rateρis the momentumDefaults:
η = 0.001ρ = 0.9
Optimisers.Adam: Adam optimizersolve(problem, Adam(η, β::Tuple))ηis the learning rateβ::Tupleis the decay of momentumsDefaults:
η = 0.001β::Tuple = (0.9, 0.999)
Optimisers.RAdam: Rectified Adam optimizersolve(problem, RAdam(η, β::Tuple))ηis the learning rateβ::Tupleis the decay of momentumsDefaults:
η = 0.001β::Tuple = (0.9, 0.999)
Optimisers.OAdam: Optimistic Adam optimizersolve(problem, OAdam(η, β::Tuple))ηis the learning rateβ::Tupleis the decay of momentumsDefaults:
η = 0.001β::Tuple = (0.5, 0.999)
Optimisers.AdaMax: AdaMax optimizersolve(problem, AdaMax(η, β::Tuple))ηis the learning rateβ::Tupleis the decay of momentumsDefaults:
η = 0.001β::Tuple = (0.9, 0.999)
Optimisers.ADAGrad: ADAGrad optimizersolve(problem, ADAGrad(η))ηis the learning rateDefaults:
η = 0.1
Optimisers.ADADelta: ADADelta optimizersolve(problem, ADADelta(ρ))ρis the gradient decay factorDefaults:
ρ = 0.9
Optimisers.AMSGrad: AMSGrad optimizersolve(problem, AMSGrad(η, β::Tuple))ηis the learning rateβ::Tupleis the decay of momentumsDefaults:
η = 0.001β::Tuple = (0.9, 0.999)
Optimisers.NAdam: Nesterov variant of the Adam optimizersolve(problem, NAdam(η, β::Tuple))ηis the learning rateβ::Tupleis the decay of momentumsDefaults:
η = 0.001β::Tuple = (0.9, 0.999)
Optimisers.AdamW: AdamW optimizersolve(problem, AdamW(η, β::Tuple))ηis the learning rateβ::Tupleis the decay of momentumsdecayis the decay to weightsDefaults:
η = 0.001β::Tuple = (0.9, 0.999)decay = 0
Optimisers.ADABelief: ADABelief variant of Adamsolve(problem, ADABelief(η, β::Tuple))ηis the learning rateβ::Tupleis the decay of momentumsDefaults:
η = 0.001β::Tuple = (0.9, 0.999)