Model Configuration and Shapes
Goal
Choose a CPU-friendly language-model configuration, calculate its important shape relationships, and explain which settings change parameter tables versus sequence-time compute.
Write the model contract before the code
A language model has several size choices that must agree with one another. If the embedding width cannot split evenly across attention heads, or the vocabulary size disagrees with the tokenizer, the failure may appear later as a confusing reshape or checkpoint error.
So before assembling the network, write the small set of numbers that define those interfaces. For example:
vocab_size = 64
block_size = 32
n_embd = 48
n_head = 4
n_layer = 2
head_dim = 12
The token embedding table has vocab_size rows; position capacity depends on block_size; every Transformer block carries width n_embd; and n_embd must divide cleanly across the simple equal-width heads.
See how C splits across attention heads
The example configuration uses C=48 channels and H=4 heads, so each head receives D=12 channels. The token positions stay the same.
B- batch size
T- token positions, at most block_size
C = 48- embedding/residual width
Separate parameter size from sequence-time work
Configuration values affect different resources in different ways.
- Increasing
vocab_sizeenlarges token embeddings and the final vocabulary projection. - Increasing
n_embdwidens many learned matrices throughout the model. - Increasing
n_layerrepeats whole Transformer blocks. - Increasing
block_sizeallows longer sequences and increases attention work over more positions, even when most learned matrix shapes stay unchanged.
For example, changing block_size from 32 to 64 does not automatically double the shape of every linear layer. But self-attention now has more query/key position pairs to consider. That is why context length can raise runtime/memory cost sharply without a matching increase in parameter count.
Before training, write a small shape table. If B=4, T=32, C=48, H=4, then each head has D=12 and the attention representation can be viewed as (4,4,32,12). Naming those dimensions first turns later reshape code into a checkable transformation instead of a guess.
Configuration values are coupled by constraints
A configuration is not a bag of independent knobs.
For multi-head attention, a common requirement is:
n_embd % n_head == 0
If n_embd=48 and n_head=4, each head can use dimension 12. Changing n_head to 5 may make the split invalid.
Likewise, block_size must be compatible with positional/data-window assumptions, and vocab_size must match the tokenizer and vocabulary head.
Write the configuration next to the checkpoint
Weights alone do not reliably reveal all of these choices.
A reproducible checkpoint package should record the configuration used to construct the model before parameters are loaded.
If someone guesses n_layer, n_head, or vocab_size incorrectly, loading may fail—or surrounding artifacts such as the tokenizer may mismatch while tensor shapes still look plausible.
Configuration is part of model identity.
Predict
Name dimensions before reshaping
Run the lab and validate one configuration. Change only one dimension at a time and predict which parameter group or computation changes. Doubling block_size, for example, does not multiply most learned weight matrices, but it increases the number of positions and attention work.
Loading lab…
Then try n_embd=50, n_head=8. A clean implementation should reject the incompatible contract before a later reshape makes the error harder to locate.