본문으로 건너뛰기
L8.2

Instruction Data: Build, Clean, and Format

Goal

By the end of this lesson, you can design instruction examples, clean them without losing provenance, split related records safely, and serialize them with a consistent training template.

A fine-tuning dataset is not just a pile of text. Each row is a behavioral example, and each cleaning rule changes what evidence the model receives.

Start with one support example:

instruction: classify the issue
input: Battery drains in two hours after the update.
target: {"issue":"battery_drain","severity":"review"}

That one target teaches vocabulary, output shape, and part of the decision rule. If similar examples disagree, training receives inconsistent supervision.

Build examples around behavior boundaries​

Easy examples matter, but boundary cases teach where one behavior stops and another begins. Include realistic variation such as different wording, missing optional information, and valid abstention cases.

If the deployed system should say “I do not have enough evidence,” include examples where abstention is the correct target. A model that sees only answerable examples can learn that it should always answer.

Keep held-out evaluation separate. Near-duplicates across train and evaluation can make memorization look like generalization.

Preserve identity before cleaning​

Suppose two rows come from the same conversation:

case-17 turn 1
case-17 turn 2

If one turn is in training and the other is in validation, the validation set contains closely related evidence. Preserve a source ID and a group key such as case_id, then split related records as a group.

Provenance also lets you trace a suspicious target back to its source after transformation.

Separate cleanup from semantic edits​

Reversible cleanup can include trimming accidental whitespace or normalizing a known field layout.

Semantic edits include changing a label, rewriting a target, merging categories, or removing examples using a quality judgment. Those edits need stronger documentation because they change the supervision itself.

Exact duplicates can overweight one pattern. Conflicting duplicates are more serious:

same input -> label A
same input -> label B

Do not silently keep both and assume optimization will decide which one is correct. Surface the conflict for review.

Formatting is part of the contract​

A chat-style example may become:

<system> classify support issue
<user> battery drains after update
<assistant> {"issue":"battery_drain"}

Training and inference should use compatible role markers, delimiters, and templates. The exact syntax depends on the model family, but template identity should be recorded with the dataset.

Split before fitting data-dependent preparation​

Some preparation steps learn values from the dataset itself. Examples include normalization statistics, filtering thresholds, label mappings, or vocabulary-like metadata.

Create the train/validation boundary first. Then fit those data-dependent choices on the training split unless the evaluation protocol explicitly requires something else. If validation data helps choose the preprocessing state, information from the evaluation set has leaked into the training pipeline.

This is the same leakage rule you saw earlier in classical ML: evaluation data may measure a pipeline, but it should not quietly help build the pipeline being measured.

Produce a data report​

Record at least:

  • source and version;
  • raw and kept row counts;
  • drop reasons;
  • duplicate/conflict policy;
  • group split rule;
  • behavior distribution by split;
  • template version;
  • known limitations.

That report makes the dataset auditable instead of treating preprocessing as an invisible step.

Predict

Two turns come from the same support case. What is the safer train/validation split?

Run the Browser Lab​

The Lab combines two checks: behavior coverage and group-safe splitting.

  1. Run the starter and inspect the label, abstention, and boundary counts.
  2. Complete group_split so every case_id is assigned wholly to training or validation.
  3. Change the requested validation group and confirm the function still keeps groups disjoint.
  4. Add a new boundary or abstention example and explain what distinct behavior it teaches.

Loading lab…

Quick Check

1. Why include boundary and abstention examples?
2. Why keep source and group IDs during preparation?
3. Why must training and inference formatting be compatible?

0 of 3 questions answered.

Explain it back​

Describe a preparation pipeline for records that contain several rows per source conversation. Include behavior coverage, provenance, conflict handling, group splitting, and template versioning.

Key Takeaways

  • Instruction examples teach behavior through task/input/target evidence.
  • Boundary and abstention cases matter when the product needs those distinctions.
  • Preserve provenance before transformations.
  • Keep related records together across train/evaluation splits.
  • Document semantic edits, duplicate policy, and template identity.

Next Lesson

Next, use the prepared examples in supervised fine-tuning and decide which token errors contribute to the training objective.

References

Lesson actions

Completion is stored locally on this device.

View progress