본문으로 건너뛰기
L1.10

Logistic Regression and Decision Boundaries

Goal

By the end of this lesson, you can interpret logistic-regression scores, explain how a threshold creates class labels, and probe a two-feature decision boundary to see why near-boundary cases can flip after small input changes.

A classifier can produce a score before it produces a label​

Suppose two emails are both classified as spam. One receives a positive-class score near 0.51 and another near 0.99. The final label is the same, but the model is treating the cases differently.

Logistic regression first forms a weighted score from the features:

z = w1*x1 + w2*x2 + ... + b

Then it passes that score through the logistic function:

p = 1 / (1 + exp(-z))

Very negative z values map near 0, z = 0 maps to 0.5, and very positive z values map near 1. Under the model assumptions, p is used as an estimated positive-class probability.

This gives one continuous chain to keep in mind:

features -> weighted score z -> logistic function -> probability-like score p -> threshold -> class

The score is still an estimate learned from data. A number that looks like a probability is not automatically well calibrated.

The threshold is a decision policy​

Suppose the model produces scores 0.20, 0.55, and 0.90.

At threshold 0.5, the labels are negative, positive, positive. At threshold 0.8, they become negative, negative, positive.

The model scores did not change. The policy that converts scores into actions changed. That distinction matters when false positives and false negatives have different costs.

A decision boundary is where the predicted class changes​

With two features, every possible input can be pictured as a point on a map. One region receives class 0 and another receives class 1. The border between those regions is the decision boundary.

For two-feature logistic regression, the 0.5 boundary is a straight line because the underlying weighted score is linear.

A point near that line can move from one predicted class to the other after a small change.

Probe scores and the boundary in the Lab​

The Lab fits logistic regression to study-hours and attendance features, then evaluates several probe points.

  1. Run the starter and inspect the weighted scores, positive-class probabilities, and labels.
  2. Find decision_threshold = 0.5. Change only 0.5 to 0.4.
  3. Before running, predict whether the weighted scores or probabilities will change. Then run and confirm that only the final thresholded labels can change.
  4. Reset. Identify the probe closest to probability 0.5.
  5. Lower only that probe's study-hours value a little and predict whether its weighted score and probability will cross the 0.5 boundary.
  6. Run again and record the before/after weighted score, probability, and class.
  7. Reset, then change only the attendance value by a small amount and compare how strongly the fitted model responds.

Loading lab…

The useful evidence is not merely “the label changed.” Record which feature changed, how much the score moved, and whether the point crossed the threshold.

Near the boundary does not mean wrong​

A near-boundary example is close to the model's current dividing line. That can make it sensitive to measurement noise or policy changes, but it does not prove the example is mislabeled.

Also, a clean-looking boundary can be poorly supported by data. Ask whether the point is represented by training data, whether an important feature is missing, whether labels are reliable there, and whether the behavior repeats across validation folds.

More flexible models can draw more complicated boundaries. That can capture real structure, but it can also make overfitting easier.

Keep the layers separate​

A useful mental model is:

features -> model score -> threshold -> predicted class -> action

Changing a threshold changes the decision rule without refitting the model. Changing model parameters changes the score. Changing the real-world action policy may happen even after a class is produced.

Quick Check

1. How does logistic regression turn its weighted score z into a 0–1 score?
2. What can change a class label without refitting logistic regression?
3. For two-feature logistic regression at threshold 0.5, where is the decision boundary?

0 of 3 questions answered.

Key Takeaways

  • Logistic regression forms a weighted score z and maps it through the logistic function into the 0–1 range.
  • At z = 0, the logistic score is 0.5; with a 0.5 threshold this connects directly to the linear decision boundary.
  • A threshold converts scores into labels and can be changed without refitting.
  • A decision boundary separates feature-space regions assigned different classes.
  • Near-boundary cases are useful sensitivity probes, not automatic mistakes.
  • A neat boundary is model geometry, not proof of trustworthy evidence.

Next Lesson

Next, you will look inside classification errors so accuracy does not hide which kind of mistake the model makes.

References

Lesson actions

Completion is stored locally on this device.

View progress