Structured Outputs
Goal
Define a small output schema, distinguish valid syntax from valid meaning, and validate model-produced structured data before downstream code trusts it.
Free-form prose is convenient for humans but awkward for software. If the next program step needs an item name, quantity, and confidence flag, a structured object is easier to validate than a paragraph.
Start with an explicit contract
Suppose the required output is:
{
"item": "battery",
"quantity": 2,
"needs_review": false
}
A simple schema can require:
itemis a non-empty string;quantityis an integer greater than or equal to 1;needs_reviewis a boolean;- no required field is missing.
Now the application has something concrete to check.
Valid JSON is not the same as valid output
This is valid JSON:
{"item": "", "quantity": -4, "needs_review": "maybe"}
But it violates the intended schema. Likewise, a schema-valid object can still contain a factual error:
{"item": "battery", "quantity": 9, "needs_review": false}
when the source document clearly says quantity 2. There are therefore at least three layers:
- syntax validity — can the data be parsed?
- schema validity — do fields and types satisfy the contract?
- semantic/grounding validity — are the values supported by the source and task rules?
Do not stop at layer 1.
Validation belongs before side effects
Imagine structured output controls a later action:
model output
→ create refund
The workflow should not perform the action merely because the text looks JSON-like. A safer path is:
model output
→ parse
→ schema validation
→ business-rule checks
→ authorization/human approval where required
→ action
The model proposes data. The application decides whether that data is acceptable for the next step.
Repair and retry need limits
If validation fails, a workflow can sometimes retry with the validation error:
quantity must be an integer >= 1
But retries should be bounded. An endlessly retrying model loop is not reliability. Record which validation failed, how many attempts were made, and what happens after the retry budget is exhausted.
Distinguish missing, null, and empty values
Structured contracts become clearer when absence has one deliberate representation. These three values can mean different things:
{}
{"answer": null}
{"answer": ""}
An omitted field may mean the producer violated the schema. A null field can deliberately mean “no supported answer.” An empty string may mean a present-but-empty value—or may simply be a malformed answer. Choose the meaning explicitly. For an evidence-grounded workflow, a strong contract might say:
supported=true → answer and source_id must be non-empty
supported=false → answer=null and source_id=null
Now downstream code does not need to guess whether an empty string means abstention.
Schemas evolve, so consumers need a compatibility rule
Suppose version 1 returns:
{"item": "battery", "quantity": 2}
and version 2 adds:
{"item": "battery", "quantity": 2, "needs_review": false}
Will old consumers ignore the new field, reject it, or require an explicit schema version? There is no single answer for every system, but silent assumptions are risky. When structured output becomes an API between components, treat schema changes like ordinary interface changes: version them, test them, and decide which changes are backward-compatible.
Structure gives software a boundary it can enforce
Structured generation is most useful when the next component needs dependable fields rather than prose that “looks about right.” Parsing answers one question: is this valid JSON? Schema validation answers a second: are the required keys and types present? Business and grounding checks answer further questions: is the value allowed, is a cited source among the supplied sources, and is the claim actually supported?
Keeping those checks separate makes failures easier to understand. A model can produce syntactically perfect JSON with a fabricated value, or a factually correct sentence that violates the expected schema. Neither should be silently accepted. Treat model output as untrusted data crossing an interface: parse it, validate its shape, validate task-specific rules, and only then allow downstream side effects.
Predict
Complete the Lab validator
The starter contains a small dictionary validator.
- Run it unchanged and observe the failing checks.
- Complete
validate_orderso it checks required fields, types, and quantity range. - Confirm that the valid case passes.
- Confirm that negative quantity and string-valued
needs_reviewfail. - Add one new invalid case and explain which contract rule it tests.
Loading lab…
Quick Check
Explain it back
Define a three-field schema for a task of your choice. Give one example that fails syntax, one that passes syntax but fails schema, and one that passes schema but is factually unsupported.
Key Takeaways
- Structured outputs make model results easier for software to inspect.
- Parsing, schema validation, and grounding are different checks.
- Validate before side effects or privileged actions.
- Hard business rules belong in application logic when possible.
- Retries should be bounded and their failures recorded.
Next Lesson
Complete the mini checkpoint, then move to context windows and the problem of relevant information being present but poorly used.
References
- JSON Schema, Draft 2020-12.
Completion is stored locally on this device.