Skip to main content
L8.2

Data for Instruction Tuning

Goal

Build instruction-tuning examples whose task, input, and target response are unambiguous, and design data that covers important behaviors rather than repeating one easy pattern.

Fine-tuning learns from examples. If the examples do not clearly express the desired behavior, more optimization cannot repair the ambiguity. A useful instruction example usually contains at least:

instruction / task
input or context
target response

One example is a behavioral demonstration​

Suppose the task is to convert a support note into a compact record. Input:

Battery drains in two hours after the update.

Target:

{"issue":"battery_drain","severity":"review","evidence":"two hours after the update"}

The target teaches several things simultaneously:

  • which label vocabulary to use;
  • which fields are required;
  • how much evidence to quote;
  • how terse the output should be.

This is why target quality matters so much. The model does not receive a separate hidden document saying “ignore inconsistent formatting.” It learns from the examples you provide.

Cover the decision boundary, not only easy cases​

If every training example is an obvious battery problem, the model may not learn how to separate:

  • battery drain from charging failure;
  • unsupported claims from supported evidence;
  • review from urgent;
  • missing fields from empty-string fields.

Include examples near the real boundary. Also include valid abstention/unknown behavior when the product needs it. A model that sees only answerable examples may learn that it should always answer.

Diversity should be relevant​

Useful variation can include:

  • wording and sentence order;
  • short and long inputs;
  • different but valid values;
  • missing optional information;
  • domain subtypes;
  • realistic noisy formatting.

Randomly changing text is not automatically useful diversity. The variation should represent differences the deployed system will actually encounter while preserving the same task contract.

Keep training examples out of evaluation​

If an evaluation prompt appears verbatim in training, a high score no longer measures generalization to a held-out case. Near-duplicates can leak too. For example:

train: Battery lasts two hours after update.
test: After the update, battery lasts about 2 hours.

A string-exact duplicate check will miss this semantic near-duplicate. Later cleaning work should preserve source IDs or grouping keys so related records can be split together.

Data quantity is not the first question​

Ten high-quality examples can reveal more about the task than ten thousand inconsistent ones. Before scaling volume, inspect:

  • target consistency;
  • boundary coverage;
  • source provenance;
  • duplicate rate;
  • class/behavior balance;
  • held-out coverage.

Data quality is part of the model design.

Training examples define behavior by repetition​

An instruction-tuning row is a claim about what response should follow a particular kind of input. If the dataset contains many examples of one easy pattern and almost no boundary cases, the optimization objective will faithfully emphasize that imbalance. Good data design therefore starts from the behavior table you want the model to learn, not from whatever logs happen to be easiest to collect.

Include examples that separate neighboring concepts, show what to do when evidence is missing, and represent variation the deployed system will actually encounter. At the same time, keep evaluation examples independent. Near-duplicate conversations split across train and test can make the model appear to generalize when it is mainly recognizing almost the same record.

Predict

A workflow must abstain when evidence is missing. What training data is important?

Run the Browser Lab​

The Browser Lab inspects a small instruction dataset.

  1. Click Run. Read {'rows': 3, 'labels': {'battery_drain': 1, 'charging_failure': 1, 'abstain': 1}, 'abstention_cases': 1, 'boundary_cases': 1}.
  2. Notice what is missing: there is no review example at all, so a model trained on this data never sees that label.
  3. Add a near-boundary review case and one more abstention case. After the row with "id": "c", add:
{"id": "d", "input": "battery swelling, device still works", "target": "review", "has_evidence": True, "boundary": True},
{"id": "e", "input": "please help", "target": "abstain", "has_evidence": False, "boundary": False},
  1. Before running, predict the new counts.
  2. Click Run. You should see 'rows': 5, a new 'review': 1 label, 'abstention_cases': 2, and 'boundary_cases': 2.
  3. Explain what behavior each new example teaches: d shows where “still works” becomes “needs a person to look,” and e shows that a vague request should get an abstention, not a guess.

Loading lab…

Quick Check

1. What does one instruction-tuning target teach?
2. Why include difficult boundary cases?
3. Why are near-duplicates across train and test dangerous?

0 of 3 questions answered.

Explain it back​

Create three examples for one instruction task: an easy case, a boundary case, and an abstention/missing-information case. State what distinct behavior each one teaches.

Key Takeaways

  • Instruction examples define desired behavior through task/input/target triples.
  • Target inconsistency becomes training inconsistency.
  • Include boundary and abstention cases when those behaviors matter.
  • Relevant diversity is better than arbitrary variation.
  • Training and evaluation must remain meaningfully separate.

Next Lesson

Next, clean and format the examples without destroying provenance or leaking related records across the split.

References

Lesson actions

Completion is stored locally on this device.

View progress