Skip to main content
L7.4

Few-Shot Examples

Goal

Use a small set of examples to demonstrate task behavior, choose examples that clarify difficult boundaries, and evaluate whether the examples actually improve the target cases.

Sometimes a task is easier to show than to describe. A few-shot prompt includes a small number of input/output examples before asking the model to handle a new input. The examples can demonstrate labels, format, tone, or a decision boundary.

Examples can define a label convention​

Suppose you want three labels:

BLOCK
REVIEW
ALLOW

The names are application-specific. The model needs to understand how you use them. You might provide:

Input: "contains a confirmed credential leak"
Output: BLOCK

Input: "contains an ambiguous external link"
Output: REVIEW

Input: "ordinary internal status update"
Output: ALLOW

Then the new case asks for one of those exact labels. The examples do more than show syntax. They demonstrate the intended mapping between evidence and labels.

Choose examples near the difficult boundary​

If every example is obvious, the model may still fail on the cases you care about. Suppose the real ambiguity is between REVIEW and ALLOW. Include at least one example near that boundary. This is similar to choosing useful training/evaluation data in earlier Levels: examples should represent the decisions the system actually needs to make.

Examples can teach the wrong thing​

A few-shot prompt can encode mistakes. If one example uses the wrong label, the model may imitate it. If all examples share an irrelevant pattern—perhaps every BLOCK example is very long—the model may overuse that pattern even though length is not the real rule. Examples can also carry hidden assumptions, bias, or stale policy.

So version and review important example sets as carefully as other test data.

More examples are not automatically better​

Every example consumes context-window space. Large example sets can crowd out the actual input or make the task harder to inspect. A better question is:

Which small set of examples provides the most useful evidence about the task boundary?

Sometimes zero-shot instructions are already sufficient. Few-shot prompting is a tool, not a required step.

Evaluate on held-out cases​

Do not choose examples and evaluate them on those same examples. Keep a fixed set of cases that are not used as demonstrations. Then compare:

prompt A: instruction only
prompt B: instruction + examples

under the same model and decoding settings. If prompt B improves the held-out cases, you have evidence that the examples help.

Examples act like a local specification​

A few-shot example does more than demonstrate wording. It shows what distinctions the task cares about: which input details matter, what output fields are expected, and where a decision boundary lies. That makes example selection a design problem. Three nearly identical easy examples can make the prompt look well specified while teaching almost nothing about ambiguous or boundary cases.

Prefer a small set that covers meaningfully different situations, including at least one case that could be confused with a neighboring label or output form. Keep separate evaluation cases that are not copied into the prompt. Otherwise the experiment measures whether the model can imitate demonstrations it has just seen rather than whether the demonstrated rule transfers to new inputs.

Predict

You need to improve REVIEW-vs-ALLOW decisions. Which few-shot example is usually most informative?

Run the Lab​

The Lab infers a small label convention from demonstrations.

The Lab labels a new case by copying the label of the demonstration that shares the most words with it.

  1. Click Run. For new case: ambiguous external update, the output is label: REVIEW, and the closest demonstration is ambiguous external link.
  2. Remove the closest example: delete the whole line ("ambiguous external link", "REVIEW"),.
  3. Before running, predict the new label.
  4. Click Run. The label becomes ALLOW, borrowed from ordinary internal update, which shares only the word update. Without a boundary example, the ambiguous case slid to a weaker match.
  5. Press Reset. Now keep all examples but change one label: ("ambiguous external link", "REVIEW") becomes ("ambiguous external link", "ALLOW").
  6. Click Run. The label becomes ALLOW, even though the matched example text is unchanged. One wrong demonstration moved the decision.
  7. Press Reset before continuing.

Loading lab…

Quick Check

1. What can few-shot examples communicate?
2. Why keep held-out evaluation cases?
3. Why can more examples hurt?

0 of 3 questions answered.

Explain it back​

Design three few-shot examples for a simple classification task. Identify which example is closest to the difficult decision boundary and name one held-out case you would use to evaluate the prompt.

Key Takeaways

  • Few-shot examples demonstrate behavior inside the current context.
  • Good examples clarify real decision boundaries, not only easy cases.
  • Incorrect or biased examples can teach the wrong pattern.
  • More examples consume context and are not automatically better.
  • Evaluate few-shot prompts on held-out cases under fixed conditions.

Next Lesson

Next, handle reasoning-heavy tasks through explicit decomposition, verifiable intermediate evidence, and checks rather than depending on a special phrase.

References

Lesson actions

Completion is stored locally on this device.

View progress