What Makes a Model a Reasoning Model
Many AI assistants now offer two kinds of model, or a switch between modes. One answers almost instantly. The other shows a "thinking" indicator, sometimes for a few seconds, sometimes for minutes, and then answers. The second kind is called a reasoning model, or a thinking model.
The difference is not that one model is smart and the other is not. It is about when the model does its work. A reasoning model spends extra effort working through a problem before it commits to an answer. This lesson explains what that means and why it helps.
What You'll Learn
- The difference between a fast model and a reasoning model
- What "thinking" actually is for a language model
- What test-time compute means, in plain language
- Why reasoning models are slower and cost more per answer
Fast answers vs worked answers
A standard language model writes its answer one small piece of text (a token) at a time, starting right away. Each new token is chosen based on everything before it. There is no separate planning phase. The answer is the thinking.
That works well for most requests: writing an email, explaining a concept, summarizing a page. It works less well for problems where the first step needs to be right before you can continue, such as a multi-step math problem, a tricky logic puzzle, or a bug that could be in several places.
A reasoning model adds a phase before the answer. It first writes out a long working space of intermediate steps: breaking the problem down, trying an approach, checking it, noticing a mistake, trying another. Only after that does it write the final answer you see.
Same basic technology, different amount of work before answering.
| Criteria | Fast model | Reasoning model |
|---|---|---|
| Starts answering | Right away | After a thinking phase |
| Intermediate steps | Only if you ask for them | Always, before the answer |
| Speed | Seconds | Seconds to minutes |
| Cost per answer | Lower | Higher, because of the extra tokens |
| Best at | Everyday writing and questions | Multi-step problems with a right answer |
Fast model
- Starts answering
- Right away
- Intermediate steps
- Only if you ask for them
- Speed
- Seconds
- Cost per answer
- Lower
- Best at
- Everyday writing and questions
Reasoning model
- Starts answering
- After a thinking phase
- Intermediate steps
- Always, before the answer
- Speed
- Seconds to minutes
- Cost per answer
- Higher, because of the extra tokens
- Best at
- Multi-step problems with a right answer
What "thinking" is for a language model
It is tempting to imagine the model sitting quietly and pondering. That is not what happens. A language model can only do one thing: produce the next token. So its "thinking" is more tokens: text it writes to itself before the answer.
This matters because each token the model writes becomes part of what it reads next. Writing out step 1 means step 2 is chosen with step 1 in view. By the time it reaches the answer, the model has built up a trail of relevant work to lean on, instead of having to jump straight to the conclusion.
So a reasoning model is, in a real sense, a model that has learned to show its work to itself. The next lesson explains why that helps so much.
Test-time compute
People who build these models talk about test-time compute. It is a simple idea with an unusual name.
- Training compute is the huge amount of computing used once, to create the model.
- Test-time compute is the computing used each time the model answers a question. "Test time" just means "when the model is being used."
For years, the main way to get a smarter model was to spend more on training: bigger models, more data. Reasoning models showed a second path: spend more computing at answer time, by letting the model think longer. For hard problems, a model that thinks longer can do much better than the same model answering instantly.
- More training computeBuild a bigger, better-trained model once
- More test-time computeLet the model think longer on each hard question
Many tools now let you choose how much thinking to allow, with settings like "low, medium, high" effort or a thinking budget. That setting is a dial for test-time compute.
Why it costs more
Thinking tokens are real tokens. The model has to generate every one of them, which takes time and computing power, just like the answer itself. A reasoning model might write thousands of thinking tokens to produce a short final answer.
That is why reasoning models:
- Take longer to respond.
- Cost more per question in paid plans and developer tools, or use up usage limits faster.
- Are not always the right choice. For a simple request, the extra work adds cost and delay without improving the answer.
Lesson 4 covers exactly when the extra thinking is worth it. If you want the basics of tokens and model size first, How LLMs Actually Work covers them.
Key Takeaways
- A reasoning model works through a problem in a thinking phase before writing its final answer.
- For a language model, thinking means writing more tokens: intermediate steps it can then read and build on.
- Test-time compute is computing spent each time the model answers. Reasoning models spend more of it on purpose.
- More thinking helps most on multi-step problems with a clear right answer.
- Thinking tokens make reasoning models slower and more expensive per answer, so they are not the best choice for everything.

