Optimizers: SGD to Adam
Goal
By the end of this lesson, you can explain what an optimizer adds beyond a gradient, compare plain SGD and Adam on the same objective, and recognize when changing optimizers would only hide a deeper bug.
The gradient gives evidence; the optimizer turns it into an update
Backpropagation gives each parameter a gradient: a local direction and sensitivity for the current loss.
An optimizer decides how to turn that gradient information into a parameter update over time.
The simplest case is plain stochastic gradient descent, or SGD:
parameter = parameter - learning_rate * gradient
SGD can work very well. It does not need to remember Adam-style moving averages; it can use the current gradient directly.
Why keep optimizer state?
Imagine one parameter receives gradients that change scale from step to step.
Adam keeps running summaries of:
- the gradient itself;
- the squared gradient.
Those summaries let it adapt effective update sizes across parameters and across time.
That can make early optimization easier on some problems, but Adam does not change the meaning of a wrong gradient. If the forward pass, loss, or gradient is broken, a more sophisticated optimizer still receives broken evidence.
Compare trajectories, not just final loss
Suppose SGD and Adam both start from the same parameter value on the same bowl-shaped objective.
A useful comparison keeps fixed:
- starting parameter;
- objective;
- number of steps;
- random data, if any.
Then inspect the sequence of parameter values and losses.
A final loss can hide whether one optimizer approached smoothly, oscillated, or took a large detour.
Run the optimizer comparison
The Lab minimizes the same simple quadratic with SGD and Adam.
- Click Run.
- Read the first steps before the final values. The minimum is at
w = 3.SGD first steps: [-4.0, -1.2, 0.48, 1.488]walks steadily toward it, whileAdam first steps: [-4.0, -3.5, -3.001, -2.505]moves about0.5per step, set by its learning rate. - Find
sgd_lr = 0.2and change only0.2to0.9. - Before running, predict whether the first SGD updates will move farther and whether they may overshoot the minimum at
3. - Click Run. Now
SGD first steps: [-4.0, 8.6, -1.48, 6.584]: the very first step jumps past3to8.6, then back below it. The final loss (0.231396) is worse than before even though each step was larger. - Press Reset afterward.
Loading lab…
If loss begins bouncing or growing, that is evidence about update size. It is not evidence that SGD as an algorithm is universally bad.
What to check before switching optimizers
When training behaves badly, inspect these boundaries first:
- Is the loss finite?
- Are the gradients finite?
- Do gradient signs make sense on a tiny case?
- Are parameter updates actually happening?
- Is the learning rate too large or too small?
- Is the data scale reasonable?
Only then is optimizer choice a well-grounded experimental variable.
A common misconception
“Adam is smarter, so it should fix unstable or NaN gradients.”
Adam can rescale updates, but it cannot make an invalid forward computation or undefined gradient valid. Numerical failures should be diagnosed at their source.
A little more technical: Adam's state
Adam tracks moving estimates often called first and second moments. Because those estimates start at zero, implementations also use bias correction in early steps.
You do not need to memorize Adam's complete equation yet. The important idea is that optimizer state affects future updates, so resetting or restoring an optimizer is different from restoring only model weights.
Quick Check
Key Takeaways
- Gradients describe local loss sensitivity; optimizers turn that evidence into updates.
- Plain SGD uses the current gradient directly; Adam keeps adaptive state.
- Compare optimizer trajectories under the same starting conditions.
- Learning rate remains important regardless of optimizer.
- Do not use a new optimizer to hide broken forward or gradient computations.
Next Lesson
Next, you will look at the state before the first optimizer step: how initial parameter values affect symmetry and signal scale.
References
- PyTorch, torch.optim.
Completion is stored locally on this device.