How Models Are Trained to Reason
A reasoning model is not a new kind of machine. It usually starts as a normal language model that has already learned language from huge amounts of text. What makes it a reasoning model is an extra stage of training that teaches it to think well before answering.
This lesson explains that extra stage at a high level: how a model can learn to reason through practice and feedback, why this works best on problems with checkable answers, and what surprising behaviors came out of it.
What You'll Learn
- The starting point: a model that already knows language
- How practice with feedback teaches a model to reason
- Why checkable problems like math and code are so useful for this
- Behaviors that emerge from the training, like self-checking
- Why reasoning models are stronger on some tasks than others
Starting point: a capable base model
Every reasoning model begins with a model that has been through normal training: predicting the next word across enormous amounts of text. That model already knows a lot. It can write, explain, and even produce a chain of thought if prompted.
What it lacks is a reliable habit of thinking long and carefully when a problem is hard. Its steps are often shallow, and it rarely notices its own mistakes. The next stage fixes that.
Learning from practice and feedback
The key technique is reinforcement learning: learning by trying, getting feedback on the result, and adjusting toward what worked. If you have taken Reinforcement Learning & RLHF Explained, this will feel familiar, but the feedback here is different.
At a high level, the loop looks like this:
- Give a problemOften math, logic, or code with a known answer
- Model thinks and answersWrites a chain of thought, then a final answer
- Check the answerRight or wrong, tests pass or fail
- Reinforce what workedWays of thinking that led to right answers become more likely
Repeat this across a very large number of problems, and the model gradually shifts toward the kinds of thinking that tend to produce correct answers. Nobody writes out the "right" reasoning for it. The model finds patterns of thought that work, because those are the ones that get rewarded.
Why checkable problems matter
This kind of training needs a reliable way to say "right" or "wrong." Some problems make that easy:
- Math problems with a known final answer.
- Code that either passes its tests or does not.
- Logic puzzles with one valid solution.
These are sometimes called problems with verifiable answers. The feedback is clear and automatic, so training can run at huge scale without people grading every attempt.
Compare that with the feedback used in RLHF, where people (or a model trained on their preferences) judge which answer is more helpful or polite. That works for style and tone, but it is slower and fuzzier. Reasoning training leans on problems where correctness can be checked directly.
Both use reinforcement learning, with very different kinds of feedback.
| Criteria | Preference feedback (RLHF) | Verifiable feedback (reasoning training) |
|---|---|---|
| Question asked | Which answer do people prefer? | Is the answer correct? |
| Who judges | People, or a model of their preferences | An automatic check: answer key or tests |
| Teaches mostly | Helpfulness, tone, safety | Careful, correct problem solving |
| Scales easily | Harder | Easier |
Preference feedback (RLHF)
- Question asked
- Which answer do people prefer?
- Who judges
- People, or a model of their preferences
- Teaches mostly
- Helpfulness, tone, safety
- Scales easily
- Harder
Verifiable feedback (reasoning training)
- Question asked
- Is the answer correct?
- Who judges
- An automatic check: answer key or tests
- Teaches mostly
- Careful, correct problem solving
- Scales easily
- Easier
Behaviors that emerge
One of the most interesting findings from reasoning training is that useful habits appear without being programmed. Because they lead to more right answers, models trained this way tend to learn to:
- Check their own work: "Let me verify that result."
- Backtrack: "That approach does not work. Let me try another."
- Break problems into parts before solving them.
- Spend longer on harder problems and less on easy ones.
Nobody wrote a rule saying "double-check your arithmetic." Thinking patterns that include checking simply produced more correct answers, so they were reinforced.
Stronger in some areas than others
The training method explains where reasoning models shine and where they help less:
- Big gains: math, logic, coding, data analysis, planning with clear constraints. These are close to the checkable problems the models practiced on.
- Smaller gains: open-ended writing, taste, tone, and opinion. There is no answer key for "the best birthday message," so reasoning training has less to work with.
- Transfer is partial. Skills learned on math and code do carry over somewhat to general problem solving, which is why reasoning models can help with careful analysis in many fields. But the benefit is usually largest near the training focus.
Smaller reasoning models are also often trained by learning from the thinking of larger ones, which is one reason reasoning features have spread quickly into faster, cheaper models.
Key Takeaways
- A reasoning model starts as a normal language model, then gets extra training to think well before answering.
- That training uses reinforcement learning: try, get feedback on the result, reinforce what worked.
- It relies heavily on verifiable problems like math and code, where right and wrong can be checked automatically.
- Habits like self-checking and backtracking emerge because they lead to more correct answers.
- Reasoning models help most on checkable, multi-step problems, and less on open-ended writing and taste.

