본문으로 건너뛰기
L6.4

Model Configuration and Shapes

Goal

Choose a CPU-friendly language-model configuration, calculate its important shape relationships, and explain which settings change parameter tables versus sequence-time compute.

Write the model contract before the code​

A language model has several size choices that must agree with one another. If the embedding width cannot split evenly across attention heads, or the vocabulary size disagrees with the tokenizer, the failure may appear later as a confusing reshape or checkpoint error.

So before assembling the network, write the small set of numbers that define those interfaces. For example:

vocab_size = 64
block_size = 32
n_embd = 48
n_head = 4
n_layer = 2
head_dim = 12

The token embedding table has vocab_size rows; position capacity depends on block_size; every Transformer block carries width n_embd; and n_embd must divide cleanly across the simple equal-width heads.

See how C splits across attention heads

The example configuration uses C=48 channels and H=4 heads, so each head receives D=12 channels. The token positions stay the same.

Residual stream: (B,T,C)
B
batch size
T
token positions, at most block_size
C = 48
embedding/residual width

Separate parameter size from sequence-time work​

Configuration values affect different resources in different ways.

  • Increasing vocab_size enlarges token embeddings and the final vocabulary projection.
  • Increasing n_embd widens many learned matrices throughout the model.
  • Increasing n_layer repeats whole Transformer blocks.
  • Increasing block_size allows longer sequences and increases attention work over more positions, even when most learned matrix shapes stay unchanged.

For example, changing block_size from 32 to 64 does not automatically double the shape of every linear layer. But self-attention now has more query/key position pairs to consider. That is why context length can raise runtime/memory cost sharply without a matching increase in parameter count.

Before training, write a small shape table. If B=4, T=32, C=48, H=4, then each head has D=12 and the attention representation can be viewed as (4,4,32,12). Naming those dimensions first turns later reshape code into a checkable transformation instead of a guess.

Configuration values are coupled by constraints​

A configuration is not a bag of independent knobs.

For multi-head attention, a common requirement is:

n_embd % n_head == 0

If n_embd=48 and n_head=4, each head can use dimension 12. Changing n_head to 5 may make the split invalid.

Likewise, block_size must be compatible with positional/data-window assumptions, and vocab_size must match the tokenizer and vocabulary head.

Write the configuration next to the checkpoint​

Weights alone do not reliably reveal all of these choices.

A reproducible checkpoint package should record the configuration used to construct the model before parameters are loaded.

If someone guesses n_layer, n_head, or vocab_size incorrectly, loading may fail—or surrounding artifacts such as the tokenizer may mismatch while tensor shapes still look plausible.

Configuration is part of model identity.

Predict

If `n_embd=60` and `n_head=6`, what is `head_dim`?

Name dimensions before reshaping​

Run the lab and validate one configuration. Change only one dimension at a time and predict which parameter group or computation changes. Doubling block_size, for example, does not multiply most learned weight matrices, but it increases the number of positions and attention work.

Loading lab…

Then try n_embd=50, n_head=8. A clean implementation should reject the incompatible contract before a later reshape makes the error harder to locate.

Quick Check

1. Why require `n_embd % n_head == 0` in this implementation?
2. What does `block_size` constrain?
3. Why save configuration beside a checkpoint?

0 of 3 questions answered.

Explain it back​

Describe (B,T,C) and (B,T,V) and explain where C changes into V in a decoder-only language model.

Key Takeaways

  • Configuration values define a model shape/resource contract.
  • Embedding width and head count must be compatible.
  • Longer context raises sequence-time compute even when many parameter shapes stay fixed.
  • Save the full configuration with every trained artifact.

Next Lesson

Next, L6.5 — Assemble the Model assembles the configured embeddings, Transformer blocks, normalization, and vocabulary head.

References

Lesson actions

Completion is stored locally on this device.

View progress