본문으로 건너뛰기
L1.18

Model Review: Build a Trustworthy Classifier

Goal

By the end of this lesson, you can review a classifier using a fair information boundary, a baseline, a reproducible pipeline, multiple metrics, and explicit failure/overfitting checks.

A high score is a result, not a trust argument​

Imagine two model-review notes.

The first says:

“Our classifier reached 94% accuracy.”

The second says:

“The classifier was trained without future or target-derived features, beat a majority baseline on the same held-out set, improved recall on the important positive class, and the result was reproduced from a recorded pipeline and split. It still fails on a small group of near-boundary examples.”

The second note is more useful because it gives a chain of evidence.

Trustworthy evaluation is not one metric. It is a sequence of questions that protect against different failure modes.

Review the experiment in a useful order​

A practical review can start with six questions:

  1. What is the prediction moment? What answer are we trying to know, and when must the prediction happen?
  2. Which features are allowed then? Remove target-derived or future-only information.
  3. What is the baseline? How well does a simple sensible strategy perform?
  4. Is the evaluation independent? Check splitting, preprocessing, and model-selection boundaries.
  5. Which metrics match the decision? Read accuracy with precision, recall, or other relevant measures.
  6. Where does the candidate fail? Inspect actual errors and the train/held-out gap.

Each question catches something different. A baseline cannot detect leakage. A test set cannot tell you whether recall is the right metric. Reproducibility cannot make a leaky feature valid.

Use the Lab as a model reviewer​

The Lab creates a deterministic binary dataset, a fixed split, a DummyClassifier baseline, and a scaled logistic-regression pipeline.

  1. Before running, write a short review checklist containing at least: allowed features, split integrity, baseline, multiple metrics, and reproducibility.
  2. Click Run.
  3. Inspect results and confirm that the candidate is compared with the baseline on the same held-out data.
  4. Confirm that the known leaky_answer field is not used as a feature.
  5. Read accuracy together with precision, recall, and the train/test gap.
  6. Identify at least one limitation or failure that a single headline score would hide.
  7. Change one model setting only. Find LogisticRegression(max_iter=500, random_state=42), inside the candidate pipeline and change it to LogisticRegression(max_iter=500, random_state=42, C=0.01),. Run again.
  8. Compare with the first run. Test accuracy falls from 0.889 to 0.815 and recall falls from 0.852 to 0.815, but the train-test gap shrinks from 0.04 to 0.019. Decide as a reviewer: is the new model genuinely better, or only different? (A smaller gap alone is not a reason to approve a model that performs worse on held-out data.)
  9. Press Reset afterward.

Loading lab…

Why the leaky feature must be rejected even if the score becomes perfect​

Imagine adding leaky_answer, a column copied from the target.

The model could achieve nearly perfect evaluation because the answer is already present in the inputs.

That is not a stronger model. It is a broken prediction problem.

A trustworthy reviewer must be willing to reject a spectacular number when the information boundary is wrong.

Baseline and held-out evaluation answer different questions​

A baseline asks:

Is the candidate better than a simple sensible alternative?

A held-out evaluation asks:

Does the fitted procedure work on examples that did not guide fitting?

You need both questions. A model can generalize to held-out data and still fail to beat a trivial strategy. Or it can beat a baseline on a contaminated test and still be untrustworthy.

Write an approval decision that stays close to the evidence​

A useful review conclusion might say:

“Approve for further testing, not deployment. On the fixed held-out set, the candidate beats the majority baseline and improves recall without using target-derived features. The train/test gap is small, but the dataset is limited and we have not yet checked subgroup performance or time-based drift.”

Notice what the statement does not say. It does not claim the classifier is universally safe or solved.

A reviewer should state the next evidence needed, such as:

  • subgroup performance;
  • time-based validation;
  • probability calibration;
  • robustness to measurement noise;
  • drift monitoring after deployment.

A common misconception​

“A checklist proves the model is trustworthy.”

No finite checklist proves behavior under every future condition. The purpose of the review is disciplined evidence, explicit limits, and a decision that can be revisited when conditions change.

Quick Check

1. What is the strongest reason to keep a simple baseline?
2. What should happen if a feature directly reveals the target?
3. What makes an ML comparison reproducible?

0 of 3 questions answered.

Key Takeaways

  • Trustworthy ML relies on a chain of evidence, not one score.
  • Define the prediction moment and keep target/future information out of features.
  • Compare with a baseline on the same independent evaluation.
  • Use metrics that match decision costs and inspect actual failures.
  • Reproducibility and explicit limitations make the review auditable.

Next Lesson

You are ready for the Level Project, Trustworthy ML: Compare Models Without Cheating. After the project, Level 2 moves from classical models to neural networks from first principles.

References

Lesson actions

Completion is stored locally on this device.

Level project unlocked: Trustworthy ML: Compare Models Without Cheating

View progress