Skip to main content
L6.3

Training Examples from Text

Goal

Construct deterministic context windows and shifted targets from a token stream while protecting a train/validation boundary before overlapping examples are created.

Turn one long stream into bounded training examples​

The model can learn from a very long text corpus, but each training example must fit inside its configured context window. So the data pipeline turns the long stream into bounded windows. Each window must contain enough source tokens to build both the inputs and their one-step-ahead targets.

A model with block_size = 4 therefore needs five consecutive source tokens to create four next-token training positions:

source: [10,11,12,13,14]
input: [10,11,12,13]
target: [11,12,13,14]

That extra fifth token supplies the final target.

See why overlap can fool evaluation​

Suppose the source tokens are:

0 1 2 3 4 5 6 7

With block_size=3, adjacent windows overlap heavily:

0 1 2 → target includes 3
1 2 3 → target includes 4
2 3 4 → target includes 5

If you create all windows first and then randomly place them into train and validation, a training window and validation window may share most of their tokens. The model is then evaluated on examples that are not independent of what it saw during updates.

Instead, split the source stream first. For example, tokens 0..5 might be training and 6..7 validation. Only then should each region produce its own valid windows.

This is the same principle you learned earlier with tabular data leakage, but sequence overlap makes the boundary easier to violate accidentally.

Block size changes how much prefix each example can use​

Suppose the token stream is:

A B C D E F

With a short block size, a training example may contain:

A B C D

and shifted targets:

B C D E

A later window can start farther into the stream.

Longer blocks let later positions train with more preceding context, but they also require more memory and attention computation.

This means block_size is not only a file-chunking setting. It defines the maximum training context available to positions inside those examples.

Respect document boundaries when they matter​

Blindly concatenating documents can create artificial transitions:

end of article A → beginning of unrelated article B

Sometimes a training pipeline deliberately uses separator tokens to mark that boundary. Other pipelines build examples within documents.

The correct design depends on the data and objective, but the choice should be explicit.

When debugging strange next-token behavior, inspect the actual token windows being sampled rather than assuming the dataset loader preserved the boundaries you intended.

Predict

With `block_size=8`, how many consecutive source tokens form one complete shifted example?

Split source regions before overlap​

Overlapping windows are useful for extracting many examples, but they create a leakage trap. If windows are created first and randomly split later, nearly identical spans can appear in both training and validation.

The safe teaching rule is:

  1. choose deterministic train and validation source-text regions;
  2. create windows independently inside each region;
  3. record the split boundary, tokenizer revision, block size, and sampling seed.

Run the lab on token IDs 0..N, change block size, and predict how many valid starting offsets remain.

Loading lab…

For length N and block size T, every-offset contiguous slicing gives N - T source slices of length T+1 when the region is large enough.

Quick Check

1. Why does one source slice contain `block_size + 1` tokens?
2. When should the held-out boundary be applied?
3. What can shared/overlapping source content across train and validation cause?

0 of 3 questions answered.

Transfer the idea​

Design an exact split for 1,000 tokens with the final 10% held out. State the first validation index and where training windows must stop.

Key Takeaways

  • A shifted T-position example requires T+1 source tokens.
  • Split source text before building overlapping windows.
  • Context length changes both learning context and compute cost.
  • Reproducibility requires explicit split and sampling rules.

Next Lesson

Next, write the model's size and shape contract before assembling or training it.

References

Lesson actions

Completion is stored locally on this device.

View progress