Chain-of-Thought Prompting
/AI/4 min read
A model has no scratchpad except the tokens it has already written, and each token gets one fixed pass through the network. Asking for the steps is asking for more computation — sixty tokens of working is sixty times as much.
Chain-of-thought prompting is asking the model to write out its reasoning before the answer, rather than producing the answer directly.
It works, and the reason it works is more mechanical than "thinking helps".
Where the computation happens
A model produces one token per forward pass through its layers. For a 32-layer model, that is 32 layers of processing — no more, whatever the question.
There is no hidden workspace. Nothing persists between tokens except the tokens themselves. So if a problem needs four intermediate quantities and the model is asked for the final number immediately, all four have to be resolved inside a single pass.
Ask for the working instead and the picture changes completely:
| tokens of working | layer passes applied to the question |
|---|---|
| 1 | 32 |
| 20 | 640 |
| 60 | 1,920 |
| 200 | 6,400 |
Sixty tokens of reasoning is 60× more computation spent on the same question. And it is not merely more — each new token can attend to every intermediate result already written, so step three is computed with steps one and two available as input rather than having to rederive them.
The written steps are the scratchpad. That is the whole mechanism.
What it looks like
A problem with a threshold in it:
A crate weighs 14 kg. Freight is ₹85 for the first 5 kg and ₹12 for each additional kilogram, then 18% tax is added. What is the total?
Asked for the number alone, a common failure is to multiply all 14 kg at the per-kilogram rate — 14 × 12 = 168, then ₹198.24 — which is wrong because the first five kilograms are covered by the base charge. The error is a single missed condition, and there was nowhere to notice it.
With the steps written:
weight above the base 14 − 5 = 9 kg
extra charge 9 × 12 = ₹108
subtotal 85 + 108 = ₹193
tax 193 × 0.18 = ₹34.74
total ₹227.74
The threshold is handled in step one, where it is visible, and every later step reads a number that is already on the page.
Two ways to ask
Zero-shot is a single instruction — "work through it step by step before answering" — and costs nothing. It is the first thing to try.
Few-shot shows two or three worked examples in the format you want. It costs prompt tokens on every request but buys control over the shape of the reasoning: which intermediate quantities appear, in what order, with what units. Where the steps need to be machine-readable, or where the model keeps skipping a check you care about, the examples are how you specify it.
Where it makes things worse
Tasks with nothing to decompose. Classifying tone, judging whether text is on-topic, picking one of four labels — these are single judgements. Asking for reasoning gives the model room to construct an argument away from a correct first instinct, and accuracy falls.
Long chains. Every step must be right for the conclusion to be right. At 96% per step, three steps is 88.5% and eight is 72.1% — and a wrong intermediate is worse than no intermediate, because it is now in the context and everything after conditions on it. The model will not revisit it; it will build on it.
Explanations that are not the reasoning. If a model settles on an answer and then produces steps, the steps are a justification rather than a derivation. They will look sound and will not be what produced the answer, which makes them misleading precisely where you wanted transparency.
What to take away
Chain of thought is not a trick for making a model more careful. It is a way of giving it somewhere to put intermediate results, and more forward passes in which to compute them.
Which tells you where it belongs: problems with genuine intermediate quantities, few enough steps that they can all be right, and where the working is worth its latency. Not on single judgements, and not as a substitute for checking whether the steps actually say anything.