Datasets as Tables
Goal
By the end of this lesson, you can read a dataset as rows of examples and columns of variables, identify candidate features and a target, and notice data-quality problems before modeling.
Start by asking what one row means
Consider this tiny student dataset:
| study_hours | attendance | passed |
|---|---|---|
| 2 | 70% | No |
| 5 | 88% | Yes |
| 4 | 82% | Yes |
If each row represents one student record, then the first row says that one example has 2 study hours, 70% attendance, and a No result for passed.
This gives us the basic table vocabulary:
- a row usually represents one example;
- a column represents one measured property;
- a feature is a column used as model input;
- the target or label is the answer we want to predict.
If our task is to predict passed, then study_hours and attendance are candidate features and passed is the target.
The word candidate matters. A column can exist in a table without being a sensible model input.
A numerically valid column can still be wrong
Suppose attendance is stored as fractions such as 0.70, 0.88, and 0.82. Those mean 70%, 88%, and 82% if the column contract says 0.0 to 1.0.
Now imagine one row contains 82 because someone stored that student's attendance as a percent instead of a fraction. The computer can still read the number. The meaning is inconsistent.
Other quiet data problems include:
0meaning “missing” in one system and a real zero in another;- duplicate rows for the same person;
- an age of
-3caused by a data-entry error; - a column collected after the prediction moment;
- two columns that use different units for the same quantity.
This is why understanding the table comes before fitting a model.
Keep the answer out of the inputs
If passed is the target, putting passed inside the feature matrix would let the model read the answer it is supposed to predict.
That can produce an impressive score for the wrong reason.
A useful question for every feature is:
Would I know this value at the moment I need to make the prediction?
If the answer is no, the column may create leakage even if it is present in the dataset.
Shape is a compact description of the table
In scikit-learn, feature data is commonly stored in X and targets in y.
If there are six examples and two features, then:
Xhas shape(6, 2)— six rows, two feature columns;yhas shape(6,)— one target value for each example.
Shape is not just Python bookkeeping. It tells you whether the number of examples and features matches the task you think you built.
Inspect the browser Lab
The Lab uses a six-row student table. attendance is stored as a fraction, so 0.70 means 70%.
- Click Run without editing anything.
- Find
features:,X shape:, andy shape:. Confirm the starter reports two features,X shape: (6, 2), andy shape: (6,). - Read the first row in
rowsand translate it into ordinary language:2.0study hours,0.70attendance (70%), andpassed = 0. - After the final existing row, add exactly this seventh row:
{"study_hours": 5.5, "attendance": 0.84, "passed": 1},
- Before running, predict the shapes. There is still the same pair of features, but now there are seven examples.
- Click Run. Confirm
X shape: (7, 2)andy shape: (7,). - Remove the added row to restore the original six-row starter.
Loading lab…
After the guided pass, add a different valid row of your own. Keep the same three fields and explain why only the number of examples—not the feature count—changes.
A useful data dictionary
Before modeling a real table, write down at least:
| Column | Meaning | Unit/category | Available when? |
|---|---|---|---|
| study_hours | time spent studying before the exam | hours | before prediction |
| attendance | attendance before the exam | fraction 0.0–1.0 | before prediction |
| passed | exam outcome | 0/1 | after the exam |
This simple habit catches many mistakes that code alone will not catch.
Quick Check
Key Takeaways
- Rows are examples and columns are variables.
- Features are model inputs; the target is the answer to predict.
- Shape summarizes examples by features.
- Data meaning, units, timing, missingness, and duplication matter before modeling.
- A target or future-only field must not quietly become an input feature.
Next Lesson
Next, you will split examples into different roles so the final evaluation does not secretly help the model learn.
References
- scikit-learn, Getting Started.
Completion is stored locally on this device.