ReAct, short for “Reason + Act,” is a pattern where a model alternates between writing a short Thought, choosing an Action, and reading an Observation, rather than jumping straight to an answer or a tool call. A written Thought can help the model plan its next step; the application executes the Action. Together, they form a loop the model can use to work through a task step by step.
The pattern was introduced in the 2022 paper “ReAct: Synergizing Reasoning and Acting in Language Models” by Yao et al. The paper showed that models often perform better on multi-step tasks when they interleave reasoning traces with actions, instead of trying to solve everything in one shot.
The basic pattern
A ReAct-style exchange typically looks like a repeating sequence of three parts:
Thought: I need to find out how many open pull requests are older than a week.
Action: search_pull_requests(status="open")
Observation: 12 pull requests found, with dates.
Thought: I need to filter these to only ones older than 7 days.
Action: filter_by_date(pull_requests, older_than_days=7)
Observation: 4 pull requests match.
Thought: I have enough information to answer.
Final Answer: There are 4 open pull requests older than a week.
Each Thought explains what the model is trying to figure out. Each Action is a tool request. Each Observation is the result the application gets back and appends to the next context. The cycle repeats until the model decides it has enough to give a final answer.
The Observation becomes part of the next prompt context, so ReAct is also a context engineering pattern: the system must decide what results, history, and constraints to feed back into the loop.
Why interleaving reasoning helps
Without a reasoning step, a model may call tools somewhat blindly — guessing at what to do next based only on the raw conversation, without stating its intent. This makes mistakes harder to catch and harder to explain.
By writing a short thought before each action, the model creates an inspectable trace around its runtime inference:
- Clarifies its own plan before committing to a tool call.
- Makes it easier to catch a wrong plan early, since the reasoning is visible.
- Can adjust its next action based on what the last observation actually showed, rather than following a rigid script.
This visible reasoning can help humans reviewing the agent’s behavior understand the stated reason for an action, not just what happened. It is not guaranteed to be a faithful account of how the model internally decided. Verify actions and results independently.
ReAct vs plain tool calling
Plain tool calling can work well for a single, obvious action — “look up the weather” needs no elaborate reasoning. ReAct becomes more valuable when a task needs several steps and the right next step depends on what was just learned.
For example, a plain tool-calling setup might call one tool and stop. A ReAct-style loop can chain several tool calls together, using each observation to inform the next thought, until the task is actually complete.
ReAct today
Modern model APIs often have native structured tool calling, and many models perform some reasoning internally rather than writing every Thought into visible text. In those systems, an app may implement the ReAct loop with structured tool calls instead of parsing Thought: and Action: strings.
The core idea stays the same: reason about the next step, act through a tool or environment, feed the observation back, and continue. The format changes; the loop does not.
Use the simplest control pattern that fits the task:
- Use a single direct tool call when the action is obvious and cheap.
- Use a fixed workflow when the sequence is known and predictability matters.
- Use plan-then-execute when the task benefits from an upfront plan and fewer mid-loop decisions.
- Use ReAct when each observation may change the next step.
If two actions are independent, a modern agent can sometimes request parallel tool calls in one turn. If the second action depends on the first observation, keep the ReAct sequence ordered.
ReAct and the agent loop
ReAct is one concrete way to implement the “decide next action” step inside an agent loop. The Thought is the reasoning, the Action is the tool call, and the Observation is the result fed back into the next iteration. Loop engineering still applies on top of this — the surrounding harness needs stopping conditions, step limits, and progress checks so a ReAct loop does not run forever.
A worked example
Imagine a support agent handling: “Why hasn’t my order arrived?”
Thought: I should look up the order using the customer's email.
Action: find_order(email="user@example.com")
Observation: Order #4521, status: "shipped", carrier: "FastShip".
Thought: I should check the shipment tracking to see where it is.
Action: track_shipment(carrier="FastShip", order_id="4521")
Observation: Package delayed at regional hub, new estimate: 3 days.
Thought: I now have the reason for the delay and a new estimate.
Final Answer: Your order is delayed at a regional hub and is now expected in about 3 days.
Each thought narrows down what information is still missing, and each action fills that gap.
Limitations
ReAct is not free. Writing out reasoning adds extra tokens and latency to every step, which increases cost. The reasoning is also not a guarantee of correctness — a model can write a plausible-sounding thought and still choose a wrong action. ReAct makes a trace visible; it does not make the trace infallible.
Observations can also be hostile. A web page, file, email, or tool result can contain prompt injection that tries to steer the next Thought or Action. Treat observations as untrusted data and use prompt engineering defenses, tool permissions, and validation.
Finally, ReAct loops need step limits, budgets, and stop rules. The same safeguards used for any agentic system still apply.
The key idea
ReAct interleaves short reasoning steps with tool actions, so a model can plan, act, observe the result, and adjust — one step at a time. It is especially useful for multi-step tasks where the right next action depends on what was just discovered, and it makes an agent’s decisions easier to follow and debug.