본문으로 건너뛰기
L1.0

Level 1 — Classical Machine Learning You Can Trust

Level 0 taught you to ask basic questions about an AI experiment: What information goes in? What answer are we trying to predict? What changed? What evidence says the result is useful?

Level 1 turns those questions into a complete machine-learning workflow.

Start with two prediction jobs​

Imagine two small systems:

  1. One predicts how many minutes a bus will be late.
  2. One predicts whether an email is spam or not spam.

Both are machine-learning problems, but the type of answer is different.

  • Predicting a number such as 7.5 minutes is called regression.
  • Predicting a category such as spam or not spam is called classification.

In both cases, training means choosing model settings from examples. The hard part is not merely making the model produce an answer. We also need to know whether the way we trained and tested it was fair.

For example, a model can appear excellent if information from the test set accidentally leaks into training. A model can also have a high accuracy while repeatedly making the one kind of mistake that matters most for the real decision.

That is why this Level is called Classical Machine Learning You Can Trust: you will learn not only how to fit a model, but how to check the evidence around it.

A few words you will meet​

You do not need to memorize these now. Each one will be taught again when it becomes useful.

TermPlain meaning in this Level
lossa number that says how wrong the model is according to one chosen rule
parameteran adjustable value inside a model
optimizationchanging parameters in a direction that improves the chosen objective
metrica measurement we use to evaluate behavior, such as accuracy or recall
data leakageletting information reach training that should have been kept separate for fair evaluation
overfittingdoing very well on familiar training examples but not working as well on new examples

A training objective is the quantity the training process directly tries to improve. A real-world goal is what people actually care about. Those two should be connected, but they are not automatically the same.

What you will learn​

By the end of the level, you should be able to:

  • organize a dataset into features and a target;
  • keep training information separate from final evaluation information;
  • fit and inspect regression and classification models;
  • explain loss, gradient descent, learning rate, and feature scaling;
  • use more than one feature without losing track of order or meaning;
  • distinguish a model's score from the final decision made from that score;
  • evaluate with accuracy, precision, recall, F1, and cross-validation;
  • recognize leakage, underfitting, and overfitting;
  • package preprocessing and models into a repeatable pipeline;
  • write a model review that uses a baseline, multiple metrics, and failure analysis.

You do not need advanced algebra before starting. When an equation becomes useful, the lesson will first connect it to a small example you can inspect.

What “understanding” means here​

A lesson is not finished just because a Lab ran or a Quick Check turned green.

For the important ideas in this Level, try to reach three levels of understanding:

  1. Explain it: describe the idea in ordinary language without copying the definition.
  2. Trace it: point to a small example and show where the idea appears in the numbers, table, or output.
  3. Transfer it: apply the same reasoning to a different small example.

Read Quick Check feedback even after a correct answer. The explanation is part of the lesson because it tells you why the answer follows from the idea.

The learning path​

Part 1 — Build a fair regression experiment​

L1.1–L1.3 connect training objectives to loss, show how a dataset is organized, and establish a fair train/test boundary.

L1.4 — Linear Regression introduces a simple model for predicting numbers from a straight-line relationship.

L1.5–L1.9 build the rest of the regression workflow: loss functions, gradient descent, learning rate, feature scaling, and multiple features.

After L1.9 — Multiple Features, complete Level 1 Mini Checkpoint — Regression Experiment Review. If you can explain the answers rather than only name the terms, you are ready to move from predicting numbers to predicting categories.

Part 2 — Build and review classifiers​

L1.10–L1.14 introduce classification, logistic regression, decision boundaries, decision-relevant metrics, and cross-validation.

L1.15 — Data Leakage shows how an apparently strong result can become invalid when information crosses a boundary it should not cross.

L1.16–L1.18 finish the workflow with overfitting, reproducible pipelines, and a complete model review.

How to use the Labs​

Every numbered lesson has a browser or notebook Lab matched to the lesson goal. When a Lab appears, the lesson should tell you:

  • which value or line to inspect;
  • exactly what to change, if anything;
  • which output to read;
  • what result to expect or reason about;
  • what the result means for the concept.

When code is not the main concept, you do not need to understand every Python line. Focus first on the named input, change, and output. Then connect the result back to the explanation.

A rule that protects almost every experiment​

When a score changes, ask more than “Did it get bigger?” Ask:

  • What changed?
  • What stayed fixed?
  • Did any evaluation information influence training?
  • Which examples or mistake types changed?
  • Is the metric appropriate for the decision?
  • Could another person repeat this comparison from the recorded settings?

That habit is the thread connecting the entire Level.

Level Project​

After L1.18 — Model Review: Build a Trustworthy Classifier, complete Trustworthy ML: Compare Models Without Cheating.

You will compare a simple baseline with a learned classifier while preserving the train/test boundary, excluding a deliberately leaky feature, reporting more than one metric, and documenting at least one failure/debug path.

A strong project submission should let another learner answer two questions from your evidence:

  1. Why is this comparison fair?
  2. Why does the conclusion follow from the measurements?
Lesson actions

Completion is stored locally on this device.

View progress