Level 0 — AI Foundations and Experimental Thinking
You do not need to know Python, algebra, or machine learning before you start this level.
Level 0 is where we build the small set of ideas that the rest of the course will reuse. Some words may already sound familiar—AI, machine learning, model, feature, label, training, prediction—but we will not assume that you know what they mean in this course.
You may already know AI through products such as ChatGPT, image generators, coding assistants, or AI agents. Those are real examples of modern AI, but they are not the whole meaning of AI.
- Artificial intelligence (AI) is the broad field of building computer systems that can perform tasks involving capabilities such as perception, language, prediction, reasoning, planning, or choosing actions.
- Machine learning (ML) is one major way to build AI systems. Instead of writing every useful rule or setting by hand, a learning process uses data or examples to adjust how a model behaves.
Not every AI system has to use machine learning, and not every AI product is just one model. Modern AI applications often combine learned models with ordinary software, rules, tools, databases, and workflows.
Where do ChatGPT, LLMs, and AI agents fit?
A useful first map is:
Artificial Intelligence (AI)
│
├── rule-based, search, and planning systems
│
└── Machine Learning (ML)
│
└── Deep Learning
│
├── vision models
├── speech models
└── language models
│
└── Large Language Models (LLMs)
This is a simplified map, not a set of sealed boxes. Real systems can combine several approaches.
A large language model (LLM) is a type of deep-learning model trained on very large amounts of text or other tokenized data. It learns to predict and generate sequences of tokens. Systems built around LLMs can produce text, answer questions, write code, summarize information, and perform many other language-related tasks.
ChatGPT is an AI application built around language models plus surrounding software and product features. The model is a central part of the system, but it is not the entire application.
An AI agent is also better thought of as a system, not as a special kind of model. An agent may combine a model—often an LLM—with things such as:
- instructions or goals;
- tools it can call;
- state or memory about what has happened;
- a loop that lets it observe results and choose another action.
A simplified picture is:
model
+ instructions
+ tools
+ state / memory
+ an action loop
↓
AI assistant or agent system
So an LLM can be one component inside an agent, while AI is the much larger field containing both of these ideas.
This course starts with tiny models on purpose. If you can clearly see how a small model uses data, makes a prediction, fails, and gets evaluated, the later ideas—neural networks, deep learning, tokenization, attention, Transformers, LLMs, and agent systems—have something solid to build on.
With that larger map in mind, start with one small example where every part is easy to see.
One example first: a spam filter
Imagine that you have a collection of old email messages. For each message, a person has already marked the correct answer:
Email A → spam
Email B → not spam
Email C → not spam
Email D → spam
Now a new email arrives. We want a computer program to estimate whether that new email is spam.
To do that, the program may look at information such as:
- words in the message;
- how many links it contains;
- whether the sender is known;
- other measurable information about the email.
This one spam-filter example contains most of the vocabulary we need. Here is the whole map before we unpack each word:
| Term | In the spam-filter example | Plain meaning |
|---|---|---|
| Input | the new email | the information given to the system for one prediction |
| Feature | number of links, known sender, certain words | one usable clue from the input |
| Label | spam or not spam on an old example | the target answer attached to an example |
| Supervised learning | learn from old emails that already have labels | learning from examples that come with target answers |
| Model | the learned prediction calculation | the reusable part that turns features into a prediction |
| Prediction | spam for the new email | the model's estimated answer |
You do not need to memorize the table. The point is to see that these words describe different parts of one process.
Input
The input is the information we give the system for one prediction.
In the spam example, the input is the new email and the information we can read from it.
Feature
A feature is one piece of input information that the model can use.
For a simple spam filter, examples of features could be:
number of links = 5
sender is known = no
contains the phrase "free prize" = yes
A real system may use many more features, but the idea is the same: features are the clues available to the model.
Label
A label is the target answer attached to an example when we already know the answer we want the system to learn from.
For an old email, the label might be:
spam
or:
not spam
The label is not another clue. It is the answer we want the system to learn to predict.
Supervised learning
Supervised learning means learning from examples that include both the input information and the target answer.
You can think of it as practicing with an answer key:
example + target answer
example + target answer
example + target answer
↓
learn a useful relationship
The word supervised does not mean a human watches every prediction. It means the learning examples come with target answers—the labels—that can be used to measure mistakes while learning.
Model
A model is the reusable prediction part of the system. It takes input features and produces an output according to settings that were learned from data or otherwise chosen for the task.
For now, use this mental picture:
features from a new email
↓
model
↓
spam / not spam prediction
A model is not the whole AI app. A spam-filter application may contain an inbox screen, networking code, user settings, and many other pieces. The model is the part that performs the learned mapping from input to prediction.
Models can be extremely simple or extremely large. A tiny model might combine only two or three numbers. A modern language model can contain billions of learned values. In Level 0 we deliberately use tiny models so you can see what goes in, what comes out, and why the result changes.
Prediction
A prediction is the answer the model produces for an input whose answer we are trying to estimate.
For the new email, the prediction might be:
spam
That prediction can still be wrong. A model is a learned tool, not an answer machine that is automatically correct.
So the whole spam example looks like this:
old labeled examples
↓
supervised learning
↓
model
↓
features from a new email
↓
prediction
You do not need to memorize this diagram. The next lessons will rebuild each part slowly with smaller examples.
The questions Level 0 teaches you to ask
By the end of this level, you should be able to look at a small AI or machine-learning system and ask:
- What information goes in?
- Which parts of that information are the features?
- What answer are we trying to predict?
- If we have labeled examples, what do the labels mean?
- What did the model learn or use to connect input to output?
- What evidence would make us trust the prediction?
- What should we check when the system is wrong?
These questions will become concrete because you will build, test, compare, and debug a tiny predictor yourself.
Two more words you will use often
An error is a mismatch between a prediction and the answer we expected. Errors are useful evidence because they show us where the system fails.
A baseline is a simple comparison. If a complicated model cannot beat a very simple strategy under the same conditions, we should question whether the extra complexity is helping.
You will meet both ideas again in dedicated lessons. The goal here is only to know what kind of thing the word refers to when you see it later.
Training examples and test examples have different jobs
Suppose a student practices using ten questions and then we want to know how well the student understands the topic.
Giving the student exactly the same ten questions again is not a very strong test. The student may remember them.
Machine learning has a similar problem. Some examples are used for training—they help choose the model's settings. Different examples are kept for testing—they help us check whether the learned pattern also works on examples that were not used to choose those settings.
We will make this distinction precise in L0.5 — Training Data and Test Data. For now, remember the reason: learning from an example and fairly testing on an example are different jobs.
What you need before starting
You only need to be able to:
- use a web browser;
- read a small table;
- compare simple numbers and categories;
- describe what changed before and after a small experiment.
When code appears in a Level 0 Lab, understanding Python is not a prerequisite for understanding the idea. The lesson will tell you exactly what to run or edit and what output to inspect. Treat the code as another way to represent the small tables and rules you are already learning about.
How the lessons work
The lessons do not all force you through the same checklist.
A concept lesson may spend most of its time explaining an idea with several examples. A visual lesson may ask you to move one control and observe what changes. An experiment lesson may ask you to make a prediction before running something. A debugging lesson may begin with a result that looks wrong and ask you to investigate it.
Activities are included when they help you understand the idea, not because every lesson needs a Predict, Build, or Break section.
When a browser Lab is useful, its instructions appear next to it. The instructions should tell you:
- where to make the change;
- what to change;
- how to run it;
- what output to look at afterward.
On the first run, a browser Lab may need a short moment to load its Python environment. After it is ready, the code runs locally in your browser.
How to know when an idea is becoming usable
Getting one answer correct is useful evidence, but it is not the finish line.
After a Lesson, try to do two things:
- explain the main idea in ordinary language without copying the definition;
- apply it to a slightly different example from the one you just studied.
Quick Checks are there to help with that process. When you submit an answer, read the explanation even if you were correct. The explanation should tell you why the answer fits the idea or which misconception to fix before moving on.
If you can explain the idea and use it in a new small example, you are in a much better position to build on it in the next Lesson.
How lesson numbers work
The small marker near the top of each page tells you where you are:
- L0.0 = the Level 0 overview you are reading now;
- L0.1 = the first lesson in Level 0;
- L0.2 = the second lesson;
- and so on.
The part before the dot is the Level. The number after the dot is the lesson's position inside that Level. When this course refers to another lesson, it should normally include both the number and the title so you do not have to guess what the number means.
Learning path
Part 1 — What kind of system are we looking at?
-
L0.1 — Meet AI: Patterns, Predictions, and Decisions
Start with inputs, patterns, predictions, and the idea that a prediction is evidence-based rather than guaranteed. -
L0.2 — Rules, Examples, and Data
Separate a hand-written rule from an example, a dataset, and a label. -
L0.3 — Rules vs Learning
Compare a setting chosen directly by a person with one chosen from labeled examples.
Part 2 — What does the data mean?
-
L0.4 — Features and Labels
Learn which parts of an example are inputs and which part is the target answer. -
L0.5 — Training Data and Test Data
Learn why we need examples for learning and different examples for a fair check.
Part 3 — Run a small experiment
- L0.6 — Your First AI Experiment
Make a prediction, run a tiny supervised-learning experiment, and compare what happened with what you expected.
After L0.6 — Your First AI Experiment, you will complete a short checkpoint before moving on.
-
L0.7 — Inputs, Outputs, and Predictions
Trace information through a predictor from input to output. -
L0.8 — Errors: When a Model Gets It Wrong
Use mistakes as evidence instead of treating them as something to hide. -
L0.9 — Changing One Thing at a Time
Learn why controlled changes make experiments easier to understand. -
L0.10 — Fair Comparisons and Baselines
Compare a predictor with a simple alternative under the same conditions.
Part 4 — Explain and improve what you built
-
L0.11 — From Experiment to Explanation
Turn results into a clear claim supported by evidence. -
L0.12 — Debug and Improve a Tiny Predictor
Find a failure, propose a cause, make one change, and test whether the evidence improves.
What you will be able to do at the end
By the end of Level 0, you should be able to:
- identify features and labels in a simple supervised-learning problem;
- explain the different roles of training data and test data;
- run a small prediction experiment and describe the result;
- compare a predictor with a baseline fairly;
- change one thing at a time and explain what effect it had;
- find at least one failure and investigate a plausible cause;
- explain why a model's prediction is not the same thing as certainty.
Level Project — Data Detective
After L0.12 — Debug and Improve a Tiny Predictor, you will complete Data Detective: Build and Explain a Tiny Predictor.
The project has two valid paths because local Python is not a Level 0 entry skill:
- Core browser path: use the L0.12 Lab, run one controlled unsuccessful threshold change, and submit an evidence report. No repository or terminal setup is required.
- Builder local path: implement the same reasoning in
projects/starters/l00/data_detective.pyand run the project validator.
Both paths are judged against the same Level 0 exit skills. Writing more code does not replace explaining the experiment.
If repository, terminal, or Python setup is new, open the Project Workbench before choosing the Builder path.
The project asks you to:
- understand the data and prediction task;
- run a tiny predictor;
- test it on examples that matter;
- compare it with a simple baseline;
- deliberately find or create a failure;
- debug one plausible cause;
- explain what the evidence does — and does not — prove.
You do not need to memorize the vocabulary before moving forward. Each term will be introduced again when it becomes useful.
Completion is stored locally on this device.