Skip to main content
L7.10

Hallucinations and Uncertainty

Goal

Recognize unsupported model claims, distinguish missing source support from low-confidence wording, and design a workflow that can abstain or request a source instead of inventing an answer.

Ask what supports the answer​

Suppose the context contains:

Product: XR-4
Weight: 1.8 kg
Battery: 10 hours

Question:

What is the warranty period?

The correct workflow may need to return:

I don't know from the supplied information.

A model might instead generate “two years” because that phrase is common for products. The sentence can sound perfectly natural while being unsupported.

Confidence language is not calibrated probability​

A model might say:

I'm 95% sure the warranty is two years.

That number should not automatically be treated as a statistically calibrated probability. The model can generate confidence language as text. If an application needs calibrated uncertainty, it requires dedicated evaluation and often task-specific methods. For many workflows, a simpler and stronger rule is:

If the required source support is absent, abstain or escalate.

Ground answers to supplied sources​

A grounded workflow can require each answer to point to the source material that supports it. For example:

{
"answer": "10 hours",
"source_id": "spec-17",
"supported": true
}

If no source supports the requested field:

{
"answer": null,
"source_id": null,
"supported": false
}

This does not guarantee truth—the source itself could be wrong—but it makes the support for the answer visible.

Separate absence from contradiction​

Two failure cases are different: Missing support: no supplied source says the warranty. Conflicting sources: one source says 1 year and another says 2 years.

Missing support may call for abstention or retrieval of another source. Conflicting sources may call for a freshness check, a source-priority rule, or an explicit statement that the conflict is unresolved. Do not collapse either state into a vague confidence score: the application can act on why the answer is uncertain.

Evaluate unsupported-claim rate​

Create cases where:

  • the answer is directly present;
  • the answer requires combining two supported facts;
  • the answer is absent;
  • sources conflict;
  • a distractor contains a plausible but wrong value.

Then measure whether the workflow answers, abstains, or flags conflict appropriately. This is more informative than asking the model a handful of trivia questions and counting only correct answers.

Predict

The supplied context does not mention the requested warranty period. What is the strongest default behavior?

Complete the Lab​

The Lab contains a tiny source-support dictionary.

  1. Run the starter and observe the unsupported question.
  2. Complete the TODO so answer_from_evidence returns a value only when the requested field exists.
  3. Return None and supported=False otherwise.
  4. Add a known field and confirm it answers.
  5. Add a missing field and confirm it abstains.
  6. Explain why the same rule is stronger than asking the model to “be more confident only when correct.”

Loading lab…

Quick Check

1. What makes a generated factual claim unsupported?
2. Why should self-reported confidence be treated carefully?
3. How should contradictory sources differ from absent evidence?

0 of 3 questions answered.

Explain it back​

Create three cases: supported answer, missing support, and conflicting sources. State what your workflow should return for each and what source information should be recorded.

Key Takeaways

  • Plausible language is not proof of factual support.
  • Missing source support should often trigger abstention or escalation.
  • Self-reported confidence is not automatically calibrated probability.
  • Grounding links outputs to visible sources or tool results.
  • Missing support and conflicting sources should be handled differently.
  • Evaluate unsupported claims explicitly, not only overall answer accuracy.

Next Lesson

Next, treat prompt injection as a trust-boundary problem when untrusted text tries to influence model behavior.

References

Lesson actions

Completion is stored locally on this device.

View progress