Skip to main content
L1.18

Real-World Tabular ML: Pipelines and Model Review

Goal

By the end of this lesson, you can build a leak-safe tabular preprocessing pipeline for numeric and categorical features, handle missing values, evaluate class imbalance with more than accuracy, and review several model families under the same split and evidence rules.

Real tables are rarely clean numeric matrices​

A realistic table may contain numeric features, categorical features, missing values, identifiers, and a target.

A model cannot safely learn from it until we define how each feature type is processed.

The important idea is not a collection of preprocessing tricks. It is one fitting boundary: every data-dependent transformation must learn only from the training portion used at that stage.

Missing values need an explicit rule​

An imputer fills missing values using a repeatable strategy.

For a numeric column, one simple choice is the training median. For a categorical column, one choice is the most frequent training category or a fixed missing category.

When the fill value depends on data, it is learned preprocessing.

Therefore:

split first -> fit imputer on training data -> reuse it on validation/test data

Computing a median from the entire dataset before the split leaks evaluation information.

Categories need a numeric representation​

A category such as plan type = basic / plus / pro is not an ordered number. Mapping them to 1 / 2 / 3 would invent spacing that the data did not give us.

One-hot encoding creates separate indicator columns instead.

Unknown categories at evaluation time need an explicit policy; a robust pipeline should not crash merely because a valid new category appears.

Different columns can need different preprocessing​

Numeric columns may need imputation and, for scale-sensitive models, standardization.

Categorical columns may need imputation and one-hot encoding.

A ColumnTransformer lets those branches happen in one fitted object before the model.

The full procedure becomes:

raw table -> column-specific preprocessing -> model

When this object is fitted inside cross-validation, each fold learns its imputers, encoders, and scalers only from that fold's training rows.

Class imbalance changes what accuracy means​

Suppose only 5% of examples belong to the positive class.

A model that always predicts the majority class gets 95% accuracy while finding none of the positives.

Class imbalance is not automatically a reason to resample or reweight. It is first an evaluation and data-understanding problem.

Ask about class frequencies, decision costs, precision, recall, the confusion matrix, split quality, and whether there are enough positive examples to support a reliable estimate.

Only then decide whether class weights, sampling, thresholds, or more data are appropriate.

Run a leak-safe mixed-type pipeline​

The Lab creates a deterministic table with numeric columns, categorical columns, missing values, and an imbalanced target.

  1. Run the starter.
  2. Inspect which columns are numeric and which are categorical.
  3. Confirm that imputation and encoding live inside the preprocessing pipeline.
  4. Compare the majority baseline and logistic-regression pipeline on the same held-out rows.
  5. Read accuracy together with precision and recall.
  6. Confirm that the transformed feature count is learned by training preprocessing.
  7. Change one model setting or preprocessing choice while keeping the split fixed.
  8. Run again and record what changed and what stayed fixed.

Loading lab…

A pipeline makes fitting order reproducible. It does not guarantee that the features are valid or that the split matches the future task.

Review multiple model families under the same rules​

You now know several candidate families:

  • a trivial baseline;
  • logistic regression;
  • a decision tree;
  • Random Forest;
  • gradient boosting.

A fair review keeps the information boundary and evaluation conditions consistent.

Do not choose a model merely because one headline metric is highest.

Ask whether every candidate used the same allowed information, whether model selection stayed away from the final test set, which metrics match decision costs, whether gains repeat across validation folds, whether there is an overfitting gap, and what complexity or explainability tradeoff comes with the gain.

Record enough to reproduce the comparison​

ItemRecord
Datasource/revision and feature definitions
Splitstrategy, proportions, seed, grouping or time rule
Preprocessingcolumn roles, imputation, encoding, scaling
Candidatemodel family and important hyperparameters
Selectionvalidation method and selection metric
Final evidenceheld-out metrics plus failure analysis

Another learner should be able to reconstruct the comparison without guessing.

A high score is a result, not a trust argument​

A trustworthy model review is a chain of evidence:

  1. define the prediction moment;
  2. reject future or target-derived features;
  3. keep final evaluation outside fitting and model selection;
  4. place learned preprocessing inside the fitting boundary;
  5. compare against a baseline;
  6. use metrics that expose important errors;
  7. inspect overfitting and concrete failures;
  8. state limitations.

Quick Check

1. Why should imputation and one-hot encoding be inside the fitted pipeline?
2. Why can accuracy be misleading on an imbalanced problem?
3. What makes a model-family comparison fair?

0 of 3 questions answered.

Key Takeaways

  • Real-world tabular ML needs explicit handling for numeric, categorical, and missing values.
  • Imputation, encoding, and scaling must respect the training boundary.
  • One-hot encoding represents categories without inventing numeric order.
  • Class imbalance requires metrics and error analysis beyond headline accuracy.
  • Baseline, logistic regression, tree, Random Forest, and boosting candidates should be compared under the same evidence rules.
  • Reproducible pipelines make a comparison auditable but cannot repair an invalid prediction problem.

Next Lesson

You are ready for the Level Project, Trustworthy ML: Compare Models Without Cheating. The project carries this lesson forward with numeric and categorical columns, intentional missing values, a minority positive class, fitted preprocessing inside every candidate pipeline, and several classical model families compared without changing the split, leakage rules, or review standard.

References

Lesson actions

Completion is stored locally on this device.

Level project unlocked: Trustworthy ML: Compare Models Without Cheating

View progress