본문으로 건너뛰기
L1.1

What Machine Learning Is Really Optimizing

Goal

By the end of this lesson, you can explain that training searches for model settings that make a chosen loss smaller, and you can distinguish that training objective from the larger real-world goal.

Training needs a scoreboard​

Imagine that you are trying to improve three predictions. The real targets are:

2, 4, 6

Two candidate rules produce:

  • Rule A: 3, 5, 7
  • Rule B: 7, 9, 11

Rule A misses each target by 1. Rule B misses each target by 5. Even before using machine-learning vocabulary, you can see that Rule A is closer.

A training algorithm needs a way to turn that idea of “closer” into a number it can measure. That number is usually called a loss.

For this lesson, we use mean squared error (MSE). It squares each miss and then averages the squared values.

For Rule A, each squared error is 1² = 1.

For Rule B, each squared error is 5² = 25.

So Rule A receives the smaller loss.

That gives training a practical target:

change the adjustable model settings so the chosen loss becomes smaller.

Those adjustable settings are called parameters. Later, a model may have millions of them.

The training procedure that decides how to change parameters is called an optimizer. The repeated process of changing parameters to improve an objective is called optimization.

For now, keep this chain in mind:

parameter change
→ prediction changes
→ error changes
→ loss changes
→ optimizer uses that evidence for the next change

The objective is not the whole real-world goal​

There is an important catch.

A model cannot directly optimize a vague instruction such as “be useful,” “be fair,” or “help every student.” Training needs something measurable.

Suppose a school prediction system minimizes average error across all students. The average can improve even while one small group of students is repeatedly predicted badly.

The training objective improved, but the product may still be failing an important requirement.

This is why we separate two ideas:

  • training objective: the measurable quantity the optimizer directly tries to improve;
  • real-world goal: the broader outcome people actually care about.

A useful objective should support the real-world goal, but the two are rarely identical.

Try the tiny loss experiment​

The Lab below computes MSE for three candidate prediction rules.

  1. Click Run without editing anything.
  2. Find MSE by rule: and note that Rule A starts with an MSE of 1.0.
  3. In candidates, find Rule A: "A": np.array([2.0, 4.0, 6.0, 8.0]).
  4. Change only its first prediction from 2.0 to 3.0. The matching first target is 3.0, so this removes one squared error while leaving the other three predictions unchanged.
  5. Before running, predict whether Rule A's MSE should become larger or smaller.
  6. Click Run. Look again at MSE by rule:; Rule A should move from 1.0 to 0.75 and remain the lowest-loss rule.
  7. Restore 3.0 to 2.0 before moving on.

Loading lab…

After this guided pass, try one additional prediction change of your own. Keep every other prediction fixed and explain the direction of the resulting MSE change before you run it.

The important result is not merely which rule wins. It is the chain of reasoning:

prediction changes → error changes → loss changes → the objective gives the optimizer a preference.

A common misconception​

“If training loss is low, the system must be good.”

Low loss is useful evidence about the objective you chose. It is not proof that the system is fair, safe, robust to new conditions, gives reliable confidence estimates, or is useful for every group.

Later lessons add held-out evaluation, multiple metrics, baselines, and failure analysis precisely because one training number cannot answer every question.

Transfer the idea​

Suppose you are predicting bus arrival times. A loss based on average arrival-time error may be sensible for training.

But a transit team might also care about rare delays, reliability on less frequent routes, or whether errors are worse during accessibility-critical trips. Those are additional evaluation questions, not automatically captured by one average loss.

Quick Check

1. What does training directly try to improve in this lesson?
2. If one prediction error grows from 2 to 6 under squared error, what happens to its contribution to loss?
3. Why can low training loss still be insufficient?

0 of 3 questions answered.

Key Takeaways

  • Training needs a measurable objective.
  • Loss summarizes prediction error according to a chosen rule.
  • Parameters are adjustable model settings.
  • An optimizer changes parameters; optimization is the repeated search for settings that improve the objective.
  • A better objective value is evidence about that objective, not proof of complete real-world success.

Next Lesson

Next, in L1.2 — Datasets as Tables, you will examine the object that feeds every classical ML experiment: a dataset organized as examples and variables.

References

Lesson actions

Completion is stored locally on this device.

View progress