Skip to main content
L7.5

Reasoning Tasks Without Magic Words

Goal

Break a difficult task into checkable stages, use ordinary code or source references where they can verify a result, and explain why no special reasoning phrase can guarantee a correct answer.

Decompose the job into observable stages​

Suppose the task is:

Choose the cheaper shipping option after tax and a fixed handling fee.

Instead of asking for one final sentence, define the computation:

1. read base prices
2. apply tax
3. add handling fee
4. compare totals
5. return selected option and totals

Now intermediate values can be checked. If the final choice is wrong, you can determine whether the error came from reading the inputs, arithmetic, or comparison. This is the same debugging habit used in forward passes and Transformer traces.

Ask for useful evidence, not private hidden reasoning​

A workflow does not need an unlimited transcript of every internal thought. What it needs is task-relevant evidence that can be validated. For arithmetic, request the calculated numbers. For document analysis, request supporting quotes or source IDs.

For classification, request the matched rule or evidence field. For a plan, request the concrete steps the application will actually execute. The goal is not “make the model talk more.” The goal is “make correctness easier to inspect.”

Verification can happen outside the model​

If the model extracts one shipping option like this:

{"base": 12.0, "tax": 1.2, "handling": 2.0}

the application can calculate 12.0 + 1.2 + 2.0 = 15.2 itself. That is often stronger than asking the model to both read the fields and do the arithmetic. Ordinary code can perform the calculation deterministically. Use the model where language interpretation is useful. Use normal software where rules or calculations can be enforced exactly.

Use checkpoints, not extra steps for their own sake​

Keep using the shipping task. A useful workflow can make the handoffs explicit:

extract each option's base price, tax, and handling fee
↓
calculate each total in code
↓
compare the totals
↓
return the cheaper option and both totals

If the model extracts the wrong handling fee for option B, the later arithmetic can be perfectly correct and the final choice can still be wrong. Splitting the task helps because the application can inspect the first incorrect handoff instead of staring only at the final answer.

A useful stage therefore produces something observable: selected numbers, source IDs, structured fields, computed values, or a decision that can be checked against a rule. Ordinary code can verify arithmetic. A schema can reject a missing field. A source-grounded workflow can check whether a cited source was actually supplied.

More stages are not automatically better. A chain of three model calls that merely says “think again” three times adds cost without creating a new check. Decomposition earns its complexity when each stage has a distinct job and a meaningful success condition.

A multi-step workflow can still fail because it starts from a wrong assumption, omits a constraint, or receives contradictory information. Evaluation should include cases such as:

  • missing values;
  • contradictory source text;
  • values exactly on a decision limit;
  • distracting but irrelevant text;
  • cases where the correct result is “insufficient information.”

Predict

Which workflow is easier to verify?

Complete the Lab​

The Lab contains a tiny decision pipeline.

  1. Click Run unchanged. The output shows totals: {'A': 0.0, 'B': 0.0} and the checks fail, because total_cost is not written yet.
  2. Complete the TODO in total_cost. The total is the base price, plus tax on the base price, plus handling.
  3. Click Run again. Check your totals by hand: A is 12.0 + 1.2 + 2.0 = 15.2, and B is 13.0 + 0.65 + 1.0 = 14.65. The output should say selected: B, and every check should pass.
  4. Now change one extracted value while keeping the decision rule (min) fixed: change B's "handling": 1.0 to "handling": 2.0.
  5. Before running, explain which intermediate value should change first, and whether the selection will flip.
  6. Click Run. Only B's total changes, to 15.65, and the selection flips to A. Because the pipeline shows its intermediate totals, you can see exactly why the decision changed.

Loading lab…

Quick Check

1. What is the main benefit of decomposing a reasoning-heavy task?
2. When should ordinary code replace an LLM subtask?
3. What kind of intermediate output is most useful?

0 of 3 questions answered.

Explain it back​

Take one multi-step task and split it into stages. Mark which stages need language-model interpretation, which can use deterministic code, and what evidence you would record at each boundary.

Key Takeaways

  • Reasoning reliability comes from task structure and verification, not a magic phrase.
  • Intermediate outputs should expose evidence that matters to the task.
  • Deterministic calculations and hard rules are often better enforced in code.
  • Decomposition improves debugging but does not guarantee correctness.
  • Evaluation should stress each important stage and boundary.

Next Lesson

Next, make outputs machine-checkable with structured fields and schema validation.

References

Lesson actions

Completion is stored locally on this device.

View progress