본문으로 건너뛰기
L1.17

Unsupervised Learning: K-Means and PCA

Goal

By the end of this lesson, you can distinguish supervised from unsupervised learning, experiment with K-means cluster count, project higher-dimensional data into two dimensions with PCA, and explain why discovered structure is not automatically a meaningful label or causal truth.

Sometimes there is no known answer column​

Most of Level 1 has used:

supervised: X + known y -> learn prediction

But many datasets arrive without a target label:

unsupervised: X only -> look for useful structure

Unsupervised learning does not mean “the computer discovers the truth.” It means the learning procedure is not given a target y to predict.

We will use two examples: K-means groups nearby points into clusters, while PCA compresses many numeric dimensions into fewer dimensions while preserving important variation.

K-means asks for a number of clusters​

K-means starts with a chosen value k. It repeatedly assigns each point to the nearest center, moves each center to the mean of its assigned points, and repeats until the result settles.

Changing k changes the question.

With k = 2, the method must describe the data using two groups. With k = 4, it must describe the same data using four.

A lower within-cluster distance for larger k is not proof that the extra groups have real-world meaning.

Cluster numbers are not labels with built-in meaning​

A result such as cluster 0, cluster 1, cluster 2 does not tell you that the groups are “beginner,” “intermediate,” and “advanced.”

Those names would be an interpretation added by people.

Before attaching meaning to clusters, inspect the features, stability, domain context, and whether the grouping helps the actual task.

PCA asks which directions carry the most variation​

Suppose every row has four numeric features. Plotting four dimensions directly is hard.

Principal Component Analysis (PCA) finds new directions through the data and orders them by how much variation they capture.

Keeping the first two components represents each row with two new coordinates.

That can help visualization or compression, but information is lost whenever the kept components explain less than 100% of the variation.

At this level, treat PCA as a geometric compression tool. You do not need eigenvector derivations.

See the same rows in two dimensions​

Same PCA projection, different K-means questions

These are the same 36 four-dimensional Lab rows projected into two PCA coordinates. Change k to see how K-means redraws the grouping without changing the PCA positions.

k = 3 · cluster sizes 12 / 12 / 12 · PCA retained variation 98.2%
Cluster 0: circleCluster 1: squareCluster 2: triangle
PCA projection with K-means cluster assignmentsThirty-six projected points keep the same two PCA coordinates while marker shape changes with the selected K-means cluster count.PCA component 1PCA component 2Point 1, cluster 2Point 2, cluster 2Point 3, cluster 2Point 4, cluster 2Point 5, cluster 2Point 6, cluster 2Point 7, cluster 2Point 8, cluster 2Point 9, cluster 2Point 10, cluster 2Point 11, cluster 2Point 12, cluster 2Point 13, cluster 0Point 14, cluster 0Point 15, cluster 0Point 16, cluster 0Point 17, cluster 0Point 18, cluster 0Point 19, cluster 0Point 20, cluster 0Point 21, cluster 0Point 22, cluster 0Point 23, cluster 0Point 24, cluster 0Point 25, cluster 1Point 26, cluster 1Point 27, cluster 1Point 28, cluster 1Point 29, cluster 1Point 30, cluster 1Point 31, cluster 1Point 32, cluster 1Point 33, cluster 1Point 34, cluster 1Point 35, cluster 1Point 36, cluster 1

The point positions stay fixed because the visual uses one two-component PCA projection of the Lab fixture. Changing k changes only the K-means grouping. Notice what k = 2 is forced to merge and what k = 4 is forced to split.

The marker shapes identify clusters without relying on color. The cluster numbers themselves are arbitrary.

Change k and project to 2D​

  1. In the visual, switch between k = 2, 3, and 4 and describe which visible groups merge or split.
  2. Run the deterministic four-dimensional Lab with k = 3.
  3. Read cluster counts and K-means inertia.
  4. Change k to 2, then to 4, keeping the data fixed, and connect the numeric output to what you saw in the visual.
  5. Reset to k = 3.
  6. Inspect the PCA output shape: two coordinates per row.
  7. Read the fraction of variation retained by the two components and compare it with the 98.2% shown in the visual.
  8. Explain what the missing fraction means.

Loading lab…

Do not compare cluster IDs across runs as though “cluster 1” has a permanent identity. The numbering is only an internal assignment.

Structure is not causation​

A cluster can appear because of measurement choices, scale, sampling, or correlated features.

A principal component can summarize variation without representing a human-interpretable concept.

Unsupervised methods can reveal useful patterns, but they do not by themselves explain why the pattern exists, whether it will persist, or whether one feature causes another.

Quick Check

1. What is the key difference between supervised and unsupervised learning here?
2. What does changing k in K-means change?
3. What does a two-component PCA projection lose when retained variance is below 100%?

0 of 3 questions answered.

Key Takeaways

  • Supervised learning uses X plus known y; unsupervised learning searches X for structure without target labels.
  • K-means groups points around k learned centers.
  • Changing k changes the grouping question and does not automatically reveal true real-world categories.
  • PCA compresses numeric dimensions into directions that preserve large amounts of variation.
  • Clusters and components are useful structures, not automatic semantic labels or causal truths.

Next Lesson

Next, you will combine numeric and categorical preprocessing, missing-value handling, imbalance-aware evaluation, pipelines, and model comparison into a realistic tabular ML workflow.

References

Lesson actions

Completion is stored locally on this device.

View progress