Python course Β· Module 10: AI and Language Models

ReAct Agents - Reasoning and Acting

13 min read
In this lesson7

Welcome! Darwin here. Ask an ordinary model: "What is the weather in Serengeti and what animals live there?". It will answer smoothly, but it will make up the weather, because it has no access to any weather station. In earlier lessons we gave the model tools. Now we will teach it to use them step by step, like a tracker who first thinks, then follows the tracks, and finally assesses what they found.

ReAct (Reasoning and Acting) is a design pattern for AI agents that combines reasoning with taking actions. The agent alternates between "thinking" (reasoning) and "doing" (acting), which leads to better results! The pattern was described by researchers from Princeton and Google in the 2022 paper "ReAct: Synergizing Reasoning and Acting in Language Models".

What is ReAct?

ReAct is a paradigm in which an AI agent:

  1. Thought - analyzes the situation, plans
  2. Action - performs an action (e.g. calls a tool)
  3. Observation - observes the result of the action
  4. Repeat - repeats the cycle until the problem is solved

The loop keeps turning until the agent decides it has enough information. The bottom of the diagram shows the log of one expedition:

1β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
2β”‚                    ReAct Loop                           β”‚
3β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
4β”‚                                                         β”‚
5β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”       β”‚
6β”‚  β”‚ Thought  │───▢│  Action  │───▢│ Observation β”‚       β”‚
7β”‚  β”‚ (Think)  β”‚    β”‚ (Act)    β”‚    β”‚ (Observe)   β”‚       β”‚
8β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜       β”‚
9β”‚       β–²                                  β”‚              β”‚
10β”‚       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜              β”‚
11β”‚                    (Repeat)                             β”‚
12β”‚                                                         β”‚
13β”‚  Example:                                               β”‚
14β”‚  Thought: I need to check the weather in Serengeti      β”‚
15β”‚  Action: get_weather("Serengeti")                       β”‚
16β”‚  Observation: Temperature 28Β°C, sunny                   β”‚
17β”‚  Thought: Now I can answer the user                     β”‚
18β”‚  Action: final_answer("The weather in Serengeti...")    β”‚
19β”‚                                                         β”‚
20β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Observation is the key: the model does not guess the result of an action, it receives the tool's real answer.

Basic ReAct implementation

The ReActAgent class stores a dict of tools, an iteration limit and the history of steps. The system prompt inserts tool descriptions, taking them from the functions' docstrings (func.__doc__), and teaches the model the Thought/Action/Answer format:

1from openai import OpenAI
2from typing import Callable
3import json
4import re
5
6client = OpenAI()
7
8class ReActAgent:
9    """An agent implementing the ReAct pattern."""
10
11    def __init__(self, tools: dict[str, Callable], max_iterations: int = 10):
12        self.tools = tools
13        self.max_iterations = max_iterations
14        self.history: list[str] = []
15
16    def _create_system_prompt(self) -> str:
17        """Creates a system prompt with tool descriptions."""
18        tool_descriptions = "\n".join([
19            f"- {name}: {func.__doc__}"
20            for name, func in self.tools.items()
21        ])
22
23        return f"""You are a helpful Safari assistant who solves problems step by step.
24
25Available tools:
26{tool_descriptions}
27
28Use this format:
29Thought: [your reasoning about what to do next]
30Action: [tool_name(arguments)]
31
32When you have enough information to answer:
33Thought: [final reasoning]
34Answer: [your answer for the user]
35
36Always think before acting. Analyze observations before taking the next step."""

The max_iterations limit is a safety fuse: without it, an agent stuck in a loop would burn tokens forever.

The parser extracts three fragments with regular expressions from the re module. The re.DOTALL flag lets the dot match newline characters too:

1    def _parse_response(self, response: str) -> tuple[str, str, str]:
2        """Parses the response into thought, action, and answer."""
3        thought_match = re.search(r'Thought:\s*(.+?)(?=Action:|Answer:|$)', response, re.DOTALL)
4        action_match = re.search(r'Action:\s*(.+?)(?=Thought:|Observation:|Answer:|$)', response, re.DOTALL)
5        answer_match = re.search(r'Answer:\s*(.+)', response, re.DOTALL)
6
7        thought = thought_match.group(1).strip() if thought_match else ""
8        action = action_match.group(1).strip() if action_match else ""
9        answer = answer_match.group(1).strip() if answer_match else ""
10
11        return thought, action, answer

A missing match raises no exception, it just returns an empty string.

Executing an action turns text like get_weather("Serengeti") into a call of a real function:

1    def _execute_action(self, action_str: str) -> str:
2        """Executes an action and returns the observation."""
3        # Parse function name and arguments
4        match = re.match(r'(\w+)\((.*)\)', action_str)
5        if not match:
6            return f"Error: Invalid action format: {action_str}"
7
8        func_name = match.group(1)
9        args_str = match.group(2)
10
11        if func_name not in self.tools:
12            return f"Error: Unknown tool: {func_name}"
13
14        try:
15            # Simple argument parsing
16            if args_str:
17                args = [arg.strip().strip('"').strip("'") for arg in args_str.split(',')]
18                result = self.tools[func_name](*args)
19            else:
20                result = self.tools[func_name]()
21            return str(result)
22        except Exception as e:
23            return f"Execution error: {str(e)}"

Every error goes back to the model as observation text instead of stopping the program. Parsing arguments with split(',') is deliberately simple and breaks on a comma inside a string - in production it is better to use native function calling, which I show below.

The run method ties everything into a loop:

1    def run(self, query: str) -> str:
2        """Runs the ReAct agent."""
3        self.history = [f"User: {query}"]
4
5        for i in range(self.max_iterations):
6            # Build context
7            context = "\n".join(self.history)
8
9            # Call LLM
10            response = client.chat.completions.create(
11                model="gpt-4.1-mini",
12                messages=[
13                    {"role": "system", "content": self._create_system_prompt()},
14                    {"role": "user", "content": context}
15                ],
16                temperature=0
17            )
18
19            llm_response = response.choices[0].message.content
20            thought, action, answer = self._parse_response(llm_response)
21
22            # Record thought
23            if thought:
24                self.history.append(f"Thought: {thought}")
25                print(f"Thought: {thought}")
26
27            # If there's a final answer
28            if answer:
29                print(f"Answer: {answer}")
30                return answer
31
32            # Execute action
33            if action:
34                self.history.append(f"Action: {action}")
35                print(f"Action: {action}")
36
37                observation = self._execute_action(action)
38                self.history.append(f"Observation: {observation}")
39                print(f"Observation: {observation}")
40
41        return "Exceeded iteration limit without finding an answer."

Every turn appends the thought, the action and the observation to history, and on the next call the model receives the whole history. temperature=0 gives decisions that are as repeatable as possible, which is why I chose gpt-4.1-mini - reasoning models such as gpt-5-mini accept only the default temperature.

ReAct usage example

The tools are ordinary functions with docstrings. Here they return hard-coded data, in a real application they would call a weather API or a species database:

1# Tool definitions
2def get_weather(location: str) -> str:
3    """Retrieves the current weather for a given location."""
4    weather_data = {
5        "Serengeti": "28Β°C, sunny, humidity 45%",
6        "Kilimanjaro": "5Β°C, cloudy, snow at the summit",
7        "Nairobi": "22Β°C, partly cloudy"
8    }
9    return weather_data.get(location, f"No weather data for {location}")
10
11def search_animals(query: str) -> str:
12    """Searches for information about Safari animals."""
13    animals = {
14        "lion": "Panthera leo - Africa's largest cat, lives in groups called prides",
15        "elephant": "Loxodonta africana - the largest land animal, lives 60-70 years",
16        "giraffe": "Giraffa camelopardalis - the tallest animal in the world, up to 5.5m tall"
17    }
18    for key, value in animals.items():
19        if key in query.lower():
20            return value
21    return f"No information found about: {query}"
22
23def calculate_distance(from_loc: str, to_loc: str) -> str:
24    """Calculates the distance between Safari locations."""
25    distances = {
26        ("Nairobi", "Serengeti"): 350,
27        ("Serengeti", "Kilimanjaro"): 280,
28        ("Nairobi", "Kilimanjaro"): 200
29    }
30    key = (from_loc, to_loc)
31    reverse_key = (to_loc, from_loc)
32
33    if key in distances:
34        return f"{distances[key]} km"
35    elif reverse_key in distances:
36        return f"{distances[reverse_key]} km"
37    return f"Unknown distance between {from_loc} and {to_loc}"

The docstring is not decoration: it goes into the system prompt, and the model decides which tool to use based on it.

Now we register the tools and send the agent into the field:

1# Create the agent
2tools = {
3    "get_weather": get_weather,
4    "search_animals": search_animals,
5    "calculate_distance": calculate_distance
6}
7
8agent = ReActAgent(tools=tools)
9
10# Run
11result = agent.run("What's the weather in Serengeti and what animals live there?")
12print(f"\nFinal answer: {result}")

For this question the agent usually performs two actions - get_weather and search_animals - before returning an Answer.

ReAct with LangChain

LangChain has a ready ReAct agent, so you do not write the loop yourself: the @tool decorator turns a function into a tool and takes the description from its docstring:

1from langchain.agents import create_react_agent, AgentExecutor
2from langchain_openai import ChatOpenAI
3from langchain.tools import tool
4from langchain import hub
5
6# Tools as LangChain tools
7@tool
8def safari_weather(location: str) -> str:
9    """Retrieves weather for a Safari location."""
10    return f"Weather in {location}: 26Β°C, sunny"
11
12@tool
13def animal_info(animal_name: str) -> str:
14    """Searches for information about a Safari animal."""
15    return f"Information about {animal_name}: A magnificent African animal!"
16
17@tool
18def safari_distance(from_location: str, to_location: str) -> str:
19    """Calculates the distance between Safari points."""
20    return f"Distance from {from_location} to {to_location}: 150 km"

Now we connect the model, a ready ReAct prompt from LangChain Hub and the tools, and AgentExecutor runs the whole loop:

1# LLM and prompt
2llm = ChatOpenAI(model="gpt-4.1-mini", temperature=0)
3prompt = hub.pull("hwchase17/react")
4
5# ReAct Agent
6tools = [safari_weather, animal_info, safari_distance]
7agent = create_react_agent(llm, tools, prompt)
8
9# Executor
10agent_executor = AgentExecutor(
11    agent=agent,
12    tools=tools,
13    verbose=True,
14    handle_parsing_errors=True,
15    max_iterations=10
16)
17
18# Run
19result = agent_executor.invoke({
20    "input": "Check the weather in Serengeti and tell me about lions"
21})
22print(result["output"])

handle_parsing_errors=True sends a format error back to the model instead of stopping, and verbose=True prints every step. Mind the versions: this code works in LangChain 0.3. In LangChain 1.0, create_react_agent, AgentExecutor and hub were moved to the langchain-classic package (you import them from langchain_classic), and the recommended successor is create_agent from langchain.agents.

ReAct with OpenAI Function Calling

Modern models have built-in tool calling, so there is no text to parse. We describe the tool with a JSON schema, and the executing function receives ready arguments:

1from openai import OpenAI
2import json
3
4client = OpenAI()
5
6# Tool definitions for OpenAI
7tools = [
8    {
9        "type": "function",
10        "function": {
11            "name": "get_safari_info",
12            "description": "Retrieves information about Safari",
13            "parameters": {
14                "type": "object",
15                "properties": {
16                    "topic": {
17                        "type": "string",
18                        "description": "Topic: weather, animals, locations"
19                    },
20                    "query": {
21                        "type": "string",
22                        "description": "Detailed query"
23                    }
24                },
25                "required": ["topic", "query"]
26            }
27        }
28    }
29]
30
31def execute_safari_tool(topic: str, query: str) -> str:
32    """Executes a Safari tool."""
33    if topic == "weather":
34        return f"Weather for {query}: 28Β°C, sunny"
35    elif topic == "animals":
36        return f"Information about {query}: A fascinating Safari animal!"
37    elif topic == "locations":
38        return f"Location {query}: A popular Safari destination"
39    return "Unknown topic"

The description field plays the same role as a docstring, and required says which arguments the model must not leave out.

The loop looks familiar, only instead of regular expressions we read message.tool_calls:

1def react_with_functions(query: str) -> str:
2    """ReAct using OpenAI function calling."""
3    messages = [
4        {
5            "role": "system",
6            "content": "You are a Safari expert. Answer step by step, using available tools."
7        },
8        {"role": "user", "content": query}
9    ]
10
11    for _ in range(5):  # Max iterations
12        response = client.chat.completions.create(
13            model="gpt-5-mini",
14            messages=messages,
15            tools=tools,
16            tool_choice="auto"
17        )
18
19        message = response.choices[0].message
20        messages.append(message)
21
22        # Check for tool calls
23        if message.tool_calls:
24            for tool_call in message.tool_calls:
25                args = json.loads(tool_call.function.arguments)
26                result = execute_safari_tool(**args)
27
28                messages.append({
29                    "role": "tool",
30                    "tool_call_id": tool_call.id,
31                    "content": result
32                })
33        else:
34            # No tool calls = final answer
35            return message.content
36
37    return "Exceeded iteration limit"
38
39# Usage
40answer = react_with_functions("What's the weather like in Serengeti?")
41print(answer)

The model's reply with tool_calls must be appended to messages before the results, and each result is linked to its call through tool_call_id. tool_choice="auto" lets the model decide whether it needs a tool. The thought is not printed as Thought here, but the action and observation cycle stays the same. For new projects I recommend exactly this version, because it is the most robust against format errors.

Advanced ReAct with memory

An agent can also assess its own progress. First a structure for a single step - @dataclass generates the constructor, and field(default_factory=datetime.now) inserts the creation time:

1from dataclasses import dataclass, field
2from typing import Optional
3from datetime import datetime
4
5@dataclass
6class ThoughtStep:
7    """An agent reasoning step."""
8    thought: str
9    action: Optional[str] = None
10    observation: Optional[str] = None
11    timestamp: datetime = field(default_factory=datetime.now)

Optional[str] means the action and the observation may be empty, for example for the final thought.

An agent with reflection asks the model every few steps to assess the route so far:

1class AdvancedReActAgent:
2    """Advanced ReAct agent with memory and reflection."""
3
4    def __init__(self, tools: dict, model: str = "gpt-5-mini"):
5        self.tools = tools
6        self.model = model
7        self.client = OpenAI()
8        self.thought_history: list[ThoughtStep] = []
9        self.long_term_memory: list[str] = []
10
11    def reflect(self) -> str:
12        """Reflects on the steps taken so far."""
13        if len(self.thought_history) < 2:
14            return ""
15
16        steps_summary = "\n".join([
17            f"- Thought: {s.thought}, Action: {s.action}, Result: {s.observation}"
18            for s in self.thought_history[-3:]
19        ])
20
21        response = self.client.chat.completions.create(
22            model=self.model,
23            messages=[
24                {
25                    "role": "system",
26                    "content": "Analyze the steps taken so far and suggest improvements."
27                },
28                {
29                    "role": "user",
30                    "content": f"Steps so far:\n{steps_summary}"
31                }
32            ],
33            max_completion_tokens=1000
34        )
35
36        return response.choices[0].message.content
37
38    def should_reflect(self) -> bool:
39        """Decides whether the agent should reflect."""
40        # Reflect every 3 steps or when the last action failed
41        if len(self.thought_history) % 3 == 0 and len(self.thought_history) > 0:
42            return True
43        if self.thought_history and "error" in (self.thought_history[-1].observation or "").lower():
44            return True
45        return False
46
47    def run(self, query: str) -> str:
48        """Runs the agent with reflection."""
49        # ... implementation similar to the basic version
50        # but with added reflection
51
52        for i in range(10):
53            if self.should_reflect():
54                reflection = self.reflect()
55                print(f"Reflection: {reflection}")
56
57            # ... rest of the ReAct logic
58
59        return "Answer"

The run method is a skeleton to complete in the exercise. In reflect I used max_completion_tokens: reasoning models do not accept the older max_tokens, and the new limit also covers hidden reasoning tokens, so do not set it too low.

ReAct patterns

The basic loop can be extended in several ways. The docstring at the top is a map of the variants, and the class below it sketches the planning variant:

1"""
2Popular patterns used in ReAct agents:
3
41. BASIC ReAct
5   Thought -> Action -> Observation -> Repeat
6
72. ReAct with Reflection
8   Thought -> Action -> Observation -> Reflect -> Repeat
9
103. ReAct with Planning
11   Plan -> Thought -> Action -> Observation -> Replan -> Repeat
12
134. ReAct with Self-Consistency
14   Multiple ReAct traces -> Vote on best answer
15
165. ReAct with Retrieval (RAG-ReAct)
17   Thought -> Retrieve -> Augment -> Action -> Observation
18"""
19
20# Example: ReAct with Planning
21class PlanningReActAgent:
22    def __init__(self, tools: dict):
23        self.tools = tools
24        self.plan: list[str] = []
25        self.current_step = 0
26
27    def create_plan(self, query: str) -> list[str]:
28        """Creates an action plan before starting."""
29        # LLM creates a step-by-step plan
30        response = client.chat.completions.create(
31            model="gpt-5-mini",
32            messages=[
33                {
34                    "role": "system",
35                    "content": "Create a step-by-step plan to solve the task. Each step on a new line."
36                },
37                {"role": "user", "content": query}
38            ]
39        )
40        plan = response.choices[0].message.content.split("\n")
41        return [step.strip() for step in plan if step.strip()]
42
43    def replan(self, remaining_steps: list[str], new_info: str) -> list[str]:
44        """Updates the plan based on new information."""
45        # LLM can modify the remaining steps
46        return remaining_steps  # or modified steps

Self-consistency runs several independent traces and picks the answer most of them point to. My advice: start with basic ReAct and add variants only when you see a concrete problem.

ReAct is a powerful pattern for AI agents. It combines reasoning with acting, which leads to more thoughtful and accurate answers. In the next module you will learn even more advanced techniques - RAG and Multi-Agent systems!

Remember: a good agent is like a good tracker - it thinks first, then checks the tracks, and answers only once it has seen them with its own eyes.

Spotted a mistake in this lesson?

Check yourself

Answer the questions from this lesson. Pick an answer to see right away whether it is correct.

  1. 1. What is the ReAct pattern in the context of AI agents?

Hands-on tasks in the game

  • Vertical ordering

    Arrange the steps of the ReAct loop in the correct order:

  • Vertical ordering

    Arrange the steps of the ReAct loop in the correct order:

  • Code editor

    Create a ReActAgent class with run() and execute_action() methods.

  • Click in order

    Click the buttons to create a ReAct agent with LangChain:

  • Vertical ordering

    Arrange the ReAct steps with reflection:

  • Code editor

    Implement an agent with tools: get_weather, search_animals, calculate_distance.

Useful articles