The direct answer
Chain-of-thought is a simple trick with a big payoff: instead of asking an AI for just the answer, you ask it to show its work first. "Think step by step" is the whole technique — and it measurably improves the model's accuracy on reasoning problems.
The name comes from a 2022 research paper by Wei and colleagues, which showed that giving a model a few examples of step-by-step reasoning — or simply prompting it to reason aloud — significantly improved performance on math, logic, and commonsense tasks. With just eight worked examples, a 540-billion-parameter model reached state-of-the-art accuracy on a benchmark of math word problems.
How it works
Why does writing steps down help? Because every word the model writes becomes part of what it can see when writing the next one. Working through a problem aloud spreads the hard part across many small, easy predictions instead of one giant leap from question to answer.
It also makes mistakes catchable. A wrong final answer with no steps is a dead end; a wrong answer with steps shows you exactly where the reasoning went off the rails — and gives the model itself a chance to notice the slip.
A simple example
As an illustration, take a small arithmetic problem: "A garage charges $3 for the first hour of parking and $2 for each additional hour. What's the total for 5 hours?"
Asked for just the answer, a model might blurt out a plausible-looking wrong number. Asked to think step by step, it writes: 5 hours means 1 first hour plus 4 additional hours; 4 × $2 = $8; $3 + $8 = $11. Each step is trivially easy — and the chain of easy steps lands on the right answer.
That pattern scales: break a hard problem into a chain of easy ones, and the whole thing gets easier.
Why it matters
Chain-of-thought is the go-to technique whenever the answer needs reasoning rather than recall: math, logic puzzles, debugging, multi-step planning. It's also a transparency win — you can audit the steps instead of trusting a bare answer.
The costs are real, though: more steps mean more tokens, which means slower and pricier responses. And there's a sharper version of the technique worth knowing — ask the model to check each step as it goes, which catches errors while they're still cheap to fix.
The common misunderstanding
The common misunderstanding: the written steps prove the model reasoned correctly, the way a person would. They don't. The steps are generated text, and a model can produce confident, coherent-sounding reasoning that is simply wrong — the same way it can hallucinate facts.
Chain-of-thought reduces errors on reasoning tasks; it doesn't eliminate them, and the steps aren't a window into some inner thought process. Treat them like a colleague's working notes: useful, checkable, and occasionally mistaken.
What changed recently
What started as a prompting trick has grown into a built-in feature. Anthropic's documentation now describes "extended thinking": developers can give the model a thinking budget — a set number of tokens to reason with before answering — so the step-by-step process happens by default rather than by request. Newer models can even decide for themselves how much thinking a question deserves.
The lesson of 2022 stuck: models do better when they work things out instead of guessing. The only thing that changed is who's asking them to — the prompt, or the product.
Try it on PlainLogic
Breaking a broken build into steps is chain-of-thought in action — try it hands-on in the Debug the Build game.