The Agent Loop: Think, Act, Observe, Repeat
Every agent, from a coding assistant to a browser agent, runs the same basic cycle. It thinks about what to do, does one thing, looks at the result, and decides again. This is called the agent loop.
Once you can see the loop, agents stop being mysterious. You can follow what an agent is doing, predict where it will struggle, and understand the "thinking" and "steps" many agent tools now show on screen. This lesson walks through the loop, then follows a real-style task from start to finish.
What You'll Learn
- The four steps of the agent loop
- How to read an agent's step-by-step trace
- How an agent decides when it is finished
- Why each extra step adds time, cost, and risk
The four steps
- 1. ThinkLook at the goal and results so far, pick the next step
- 2. ActCall one tool, or write the final answer
- 3. ObserveRead what the tool returned
- 4. RepeatGo back to Think with the new information
Think. The model reads everything it has so far: the goal, the list of tools, and the results of earlier steps. It decides the single best next move. Many products show this as "thinking" or "reasoning" text.
Act. The model asks for one tool with specific inputs, such as "search the web for X" or "open this page." The software around the model runs that tool. The model itself never touches the web or your files.
Observe. The tool's result, such as search results, page text, or an error, is added to what the model can see.
Repeat. Back to Think. The loop continues until the model decides the goal is done, a person stops it, or it hits a limit on steps or cost.
This pattern is sometimes called ReAct, short for "reason and act," after the research that made it popular. You do not need the name, but you will see it in articles about agents.
Follow a real trace
Here is a simplified trace of an agent given this goal: "Find three free online Python courses for beginners that offer a certificate, and put them in a table with the course length."
Step 1
- Think: I need to find free beginner Python courses with certificates. I will search first.
- Act: web search for "free beginner Python course with certificate"
- Observe: 10 results, including course listings and review articles.
Step 2
- Think: A review article may list several options at once. I will open the most relevant one.
- Act: open the page of the top review article
- Observe: The article lists six courses, but it is two years old.
Step 3
- Think: The article may be out of date. I should check each course's own page for price and certificate details.
- Act: open the first course's page
- Observe: Free to take, but the certificate costs extra.
Step 4
- Think: That does not meet the "free certificate" need. I will skip it and check the next one.
- Act: open the second course's page
- Observe: Free, certificate included, about 10 hours.
Steps 5 to 8: The agent checks more course pages until it has three that match.
Step 9
- Think: I have three courses that meet every condition. The goal is done.
- Act: write the final answer as a table.
Notice three things in this trace:
- The agent changed its plan when it learned something new, first when the article looked old, then when a certificate turned out to cost money. That flexibility is the whole point of an agent.
- Each step depended on the step before. If step 2 had opened the wrong page, everything after would have been built on bad information.
- It took nine steps for a task a person might describe in one sentence.
When an agent tool shows you its steps, read them the same way. Look at what it searched, which pages it opened, and where it changed course. That is often the fastest way to spot a mistake.
How does an agent know it is done?
The model decides. At each Think step, it can choose to stop and give a final answer. This works well when the goal is clear, like "three courses, free certificate, in a table."
It works badly when the goal is vague. "Research the best Python courses" has no clear finish line, so an agent may stop too early with a thin answer, or keep searching far longer than needed. Clear goals with clear finish conditions make agents work better. Say what "done" looks like: how many items, which format, what conditions must be true.
Most agent products also set hard limits, such as a maximum number of steps, a time limit, or a spending cap, so a confused agent cannot run forever.
Every step has a price
Each trip around the loop means another call to the model, plus a tool call. That adds up:
- Time. Nine steps take much longer than one chat reply.
- Cost. Each step uses tokens. The model re-reads the growing history every time, so later steps cost more than early ones.
- Risk. Each step is another chance to make a mistake, and mistakes carry forward. Lesson 5 covers this.
Agents trade speed and cost for the ability to act and check.
| Criteria | One-shot chat reply | Nine-step agent task |
|---|---|---|
| Model calls | One | Nine or more |
| Time | Seconds | Often minutes |
| Can check live sources | Only if it has a search tool | Yes, many times |
| Chances to go wrong | One | One per step, and they add up |
One-shot chat reply
- Model calls
- One
- Time
- Seconds
- Can check live sources
- Only if it has a search tool
- Chances to go wrong
- One
Nine-step agent task
- Model calls
- Nine or more
- Time
- Often minutes
- Can check live sources
- Yes, many times
- Chances to go wrong
- One per step, and they add up
This is why a good rule is: use an agent when the task needs several actions or live information. For a question the model can answer from what it knows, a normal chat is faster and cheaper.
Key Takeaways
- Every agent runs the same loop: think, act, observe, repeat.
- The model only decides. The software around it runs the tools and returns results.
- Reading an agent's step-by-step trace shows where it changed plans and where mistakes started.
- Agents stop when the model judges the goal is done, so clear goals with a clear finish line work best.
- Every loop adds time, cost, and risk. Use an agent when a task needs several actions, not for simple questions.

