Skip to main content
L8.3

Cleaning and Formatting Training Data

Goal

Clean instruction data with reproducible rules, preserve provenance, prevent related-record leakage, and format examples consistently for the model input contract.

Training data usually begins as product records, documents, annotations, or logs. The model cannot safely learn from that raw collection without a preparation policy. Cleaning is not “make the data look nicer.” It is a sequence of decisions that changes what the model can learn.

Preserve raw identity before transforming​

Suppose two records come from the same support conversation:

case-17 turn 1
case-17 turn 2

If you split individual rows randomly, turn 1 can land in training and turn 2 in validation. That creates related-record leakage. Preserve a group key such as case_id and split by group when records are dependent. Also keep a source identifier so a suspicious target can be traced back to where it came from.

Separate reversible cleanup from semantic edits​

Reversible cleanup may include:

  • normalizing line endings;
  • trimming accidental surrounding whitespace;
  • converting one known field layout to another.

Semantic edits include:

  • rewriting a target;
  • dropping a label;
  • merging categories;
  • removing “low-quality” examples according to a judgment rule.

Semantic edits need stronger documentation because they change the task evidence.

Deduplicate before trusting counts​

Exact duplicate:

A → B
A → B

Conflicting duplicate:

A → B
A → C

The first may overweight one behavior. The second reveals annotation inconsistency. Do not silently keep both and assume the optimizer will discover which label is correct. Use a conflict report.

Formatting is part of the training contract​

A chat-style example might be serialized as:

<system> classify support issue
<user> battery drains after update
<assistant> {"issue":"battery_drain"}

Whatever template you use during training must agree with the tokenizer/model interface used later. If training includes role markers but inference omits them, the model sees a different input pattern. The exact syntax varies by model family. The invariant is consistency.

Split before data-dependent fitting​

If the pipeline learns anything from the corpus—label mappings, normalization statistics, filtering thresholds, vocabulary-like metadata—fit that state on training data only unless the method explicitly requires another protocol. This is the same leakage principle from earlier ML Levels.

Produce a data report​

At minimum, record:

  • source/version;
  • raw row count;
  • kept/dropped counts and reasons;
  • exact/near duplicate policy;
  • group split rule;
  • label/behavior distribution by split;
  • formatting template/version;
  • known limitations.

A fine-tuning run without a data identity is difficult to reproduce and audit.

Cleaning can change meaning, so record what was changed​

Removing exact duplicates is different from rewriting a target label or merging two conflicting examples. The first operation often removes redundant evidence; the second changes the supervision itself. Keep source IDs and case/group IDs long enough to trace those decisions and to prevent related records from crossing the train/evaluation boundary.

Formatting also becomes part of the learned pattern. If training examples use one role template or delimiter scheme but inference uses another, the model is being asked to operate under a different serialization than the one optimized during adaptation. Treat the template version as part of the dataset identity, and test formatting changes separately from content changes.

Predict

Two turns from the same support case are strongly related. What is the safer split unit?

Complete the Browser Lab cleaner​

The Browser Lab contains a small dataset with duplicate and grouped records.

  1. Click Run once. The output shows train groups: ['A', 'B'] and validation groups: ['C'], which looks correct. But one Lab check fails. The starter split the list by position (rows[:3], rows[3:]), and that only works by accident for group C. The failing check asks for {"A"} as the validation group instead.
  2. Complete the TODO in group_split: put every row whose case_id is in validation_groups into validation, and every other row into training.
  3. Click Run again and confirm every check passes.
  4. Change the call to group_split(rows, {"A", "B"}). Before running, predict both group lists.
  5. Click Run. Training is now ['C'] and validation is ['A', 'B']. No case appears on both sides, whatever groups you choose.

Loading lab…

Quick Check

1. Why keep source/group IDs during cleaning?
2. What does a conflicting duplicate indicate?
3. Why must training and inference formatting agree?

0 of 3 questions answered.

Explain it back​

Describe a cleaning pipeline for a dataset with multiple rows per source document. Include provenance, duplicate handling, group splitting, and formatting versioning.

Key Takeaways

  • Preserve source identity before transformation.
  • Split related records by group, not blindly by row.
  • Duplicate and conflict handling should be explicit.
  • Semantic edits require documentation.
  • Training/inference serialization must be compatible.
  • A data report is part of the training artifact.

Next Lesson

Next, use the prepared examples in supervised fine-tuning and identify exactly which tokens contribute to the training objective.

References

Lesson actions

Completion is stored locally on this device.

View progress