Skip to main content
L0.4

Features and Labels

Goal

By the end of this lesson, you can identify features and labels in a small dataset, explain how they play different roles, and notice when an input would not make sense for the prediction task.

A prediction problem has information and an answer​

Imagine that students take a short readiness check before moving to the next topic. After the check, each past student has been marked either Ready or Not ready.

For past students we therefore have a table like this:

Practice minutesSleep hoursReadiness label
208Not ready
458Ready
607Ready

Now imagine a new student has not taken the readiness check yet. We want to estimate whether that student will be marked Ready using information we already know beforehand.

So our prediction question is:

Can we predict the readiness label from information we know before the check?

The three columns are in the same table, but they have different roles in that question.

  • Practice minutes is information we know before the check. It can be an input clue.
  • Sleep hours is also information we know before the check. It can be another input clue.
  • Readiness label is different: it is the answer we are trying to predict for the new student.

That is the feature-versus-label distinction.

Practice minutes and Sleep hours are possible features. They are pieces of information we could give to a predictor.

Readiness label is the label. For past examples we already know it, so the model can learn from those examples. For the new student we do not know it yet—that is exactly why we need a prediction.

A useful beginner definition is:

  • Feature: information given to the model as input.
  • Label: the target answer we want the model to predict.

You will also see the word target used for the label. In many supervised-learning problems, label and target mean almost the same thing.

The same column can have a different role in a different task​

A column is not permanently a feature or permanently a label. Its role depends on the question.

Suppose we changed the question to:

Can we predict how long a student practiced from sleep hours and the recorded readiness label?

Now Practice minutes would be the label, because that is what we are trying to predict.

That is why it is better to start with a clear prediction question than to start by saying, “Here is a table. Train a model.”

The question tells us what answer we want and which information may be available before that answer is known.

A feature must make sense at prediction time​

There is an easy mistake hiding in the first table.

If we are predicting the Readiness label, it would be cheating to give the predictor that same readiness label as an input. The answer would already be inside the information used to predict the answer.

This kind of mistake is one form of data leakage.

A practical question to ask about every proposed feature is:

Will I really know this value at the moment I need the prediction?

For example:

  • Practice minutes before the readiness check: available.
  • Sleep hours from the previous night: available.
  • The readiness label produced by the check: not available yet if that is what we are predicting.
  • A number calculated using tomorrow's information: not available for a prediction made today.

A column can look useful in a completed dataset and still be an invalid feature for the real prediction task.

Not every available feature is useful​

Being available is necessary, but it does not guarantee that a feature helps.

Imagine adding a column called Student row number:

Row numberPractice minutesReadiness label
10120Not ready
10245Ready
10360Ready

The row number is easy to read and unique, but there is no reason yet to think that 101, 102, or 103 tells us anything useful about readiness.

A model can receive a feature even when that feature is useless. Later, you will learn ways to measure whether features actually help.

A small technical bridge​

In machine learning, one example is often written as:

x -> y

where:

  • x means the input feature or features;
  • y means the target label.

You do not need to memorize the symbols yet. The important idea is the separation:

information the model may use -> answer the model should predict

Try the browser Lab​

The Lab below uses the same readiness idea. You do not need to understand every Python line. Focus on three named values near the top.

  1. Click Run without changing anything.
  2. In stdout, find these lines:
    • feature: practice_minutes
    • label: readiness_label
  3. Look at the three prediction lines. With practice_minutes and threshold 40, the starter rule matches all three readiness labels.
  4. In the editor, find:
feature_name = "practice_minutes"
  1. Change only "practice_minutes" to "sleep_hours".
  2. Click Run again.
  3. Look at the predictions. The threshold is still 40, but sleep hours are values such as 7 and 8. The old threshold no longer has a sensible meaning for the new feature.

Loading lab…

This is an important result. Changing the feature did not merely change a column name. It changed what the numbers mean. A setting that made sense for practice minutes may make no sense for sleep hours.

If you want one extra experiment, change threshold = 40 to threshold = 8 while keeping feature_name = "sleep_hours", then Run again. Now the threshold is at least in the same numerical range as the feature. Notice that this does not guarantee good predictions. A feature can be valid and still be weak for the task.

A common misconception​

“The label is the answer, so giving it to the model as an input should make the model better.”

It would make the score look better, but it would destroy the prediction problem. At the moment of a real prediction, the label is exactly the thing we do not know yet.

Good machine learning is not about making a completed training table easy to copy. It is about building a procedure that can work when the target answer is still unknown.

Quick Check

1. If the task is to predict whether a new student will be marked Ready, what is the readiness-label column?
2. Why is the actual readiness label an invalid input feature for predicting that same label?
3. A feature is available at prediction time. Does that guarantee it is useful?

0 of 3 questions answered.

Key Takeaways

  • Features are the information a predictor receives.
  • The label is the target answer it is trying to predict.
  • Past examples can contain labels even though the label is unknown for the new case we want to predict.
  • A column's role depends on the prediction question.
  • A feature should be available when the real prediction is made.
  • Available does not automatically mean useful.
  • Model settings have meaning in relation to the features they operate on.

Next Lesson

Now that you can separate inputs from target answers, the next question is how to separate the examples used to choose a predictor from the examples used to check it fairly. That is the job of training data and test data.

References

Lesson actions

Completion is stored locally on this device.

View progress