ReAct 框架:思考、行动、观察
了解 ReAct 循环:模型生成 Thought,选择 Action,接收 Observation,并不断重复,直到得出 Final Answer。
ReAct 框架:思考、行动、观察 是 CoddyKit 上的免费 AI Engineering Academy 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 AI Engineering Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 AI Engineering Academy 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
What Is the ReAct Framework?
ReAct (Reasoning + Acting) is a prompting framework that interleaves the model's internal reasoning with external actions. Instead of producing a final answer immediately, the model generates a Thought, chooses an Action, receives an Observation from executing that action, and repeats the cycle until it can give a Final Answer.
The Think-Act-Observe Loop
Each iteration of the ReAct loop has three phases:
- Thought: The model reasons about what it knows and what it needs to find out next.
- Action: The model calls a tool — such as a web search, a calculator, or a database lookup.
- Observation: The result of the tool call is returned and appended to the model's context.
The loop continues until the model emits Final Answer: with its conclusion.
ReAct Trace Example
Here is a concrete trace of a ReAct agent answering 'What is the population of Tokyo?' using a search tool. Notice how each step builds on the previous observation.
# ReAct trace (pseudocode showing the reasoning loop)
# Iteration 1
Thought_1 = 'I need to find the current population of Tokyo. I will search for it.'
Action_1 = 'search("Tokyo population 2024")'
Observation_1 = 'According to the 2023 census, Tokyo has approximately 13.96 million people in the city proper.'
# Iteration 2
Thought_2 = 'I have the population. I can now give a final answer.'
Final_Answer = 'The population of Tokyo is approximately 13.96 million people.'Why ReAct Outperforms Direct Prompting
Direct prompting asks the model to answer from memory, leading to hallucinations on factual questions. Chain-of-thought prompting improves reasoning but still cannot access external information. ReAct combines both: explicit reasoning traces like chain-of-thought plus the ability to gather fresh information through tool calls.
ReAct Prompt Structure
The ReAct system prompt teaches the model the exact format to use. It lists available tools with their descriptions and shows examples of the Thought/Action/Observation/Final Answer pattern. The model learns to follow this format precisely from the prompt.
REACT_SYSTEM_PROMPT = '''You are an AI assistant with access to these tools:
- search(query: str): Search the web for up-to-date information.
- calculator(expression: str): Evaluate a mathematical expression.
- lookup(topic: str): Look up a topic in the knowledge base.
Always follow this exact format:
Thought: [your reasoning about what to do next]
Action: tool_name(arguments)
Observation: [result of the action, provided by the system]
... (repeat as needed)
Final Answer: [your final response to the user]
Begin!'''Parsing Thought, Action, and Observation
Your application code acts as the runtime for the ReAct loop. After each model response, parse the output to extract the action name and arguments, execute the corresponding Python function, format the result as an observation, append it to the message history, and call the model again.
import re
def parse_react_output(text: str):
'''Extract action name and argument from a ReAct model output.'''
action_match = re.search(r'Action:\s*(\w+)\((.*)\)', text)
if action_match:
tool_name = action_match.group(1)
tool_input = action_match.group(2).strip('\"\' ')
return tool_name, tool_input
if 'Final Answer:' in text:
answer = text.split('Final Answer:')[-1].strip()
return 'final', answer
return None, NoneThe Agent Execution Loop
The execution loop calls the LLM, parses its output, dispatches the tool, appends the observation, and loops. A max_iterations guard prevents infinite loops when the model cannot resolve a task.
from openai import OpenAI
client = OpenAI()
def run_react_agent(user_query: str, tools: dict, max_iterations: int = 10) -> str:
messages = [
{'role': 'system', 'content': REACT_SYSTEM_PROMPT},
{'role': 'user', 'content': user_query}
]
for i in range(max_iterations):
resp = client.chat.completions.create(model='gpt-4o', messages=messages)
output = resp.choices[0].message.content
messages.append({'role': 'assistant', 'content': output})
tool_name, tool_input = parse_react_output(output)
if tool_name == 'final':
return tool_input
if tool_name and tool_name in tools:
observation = tools[tool_name](tool_input)
messages.append({'role': 'user', 'content': f'Observation: {observation}'})
else:
messages.append({'role': 'user', 'content': 'Observation: Tool not found.'})
return 'Max iterations reached without a final answer.'Registering Simple Tools
Tools are just Python functions. You register them in a dictionary mapping tool name to function. The functions receive a string argument and return a string result — this keeps the interface simple and consistent with what the model expects to see as observations.
import math
def calculator_tool(expression: str) -> str:
try:
# Restrict to safe math expressions
allowed_names = {k: v for k, v in math.__dict__.items() if not k.startswith('_')}
result = eval(expression, {'__builtins__': {}}, allowed_names)
return str(result)
except Exception as e:
return f'Error: {e}'
def search_tool(query: str) -> str:
# Stub — replace with real web search API
return f'Top result for "{query}": [placeholder result]'
# Register tools
tools = {
'calculator': calculator_tool,
'search': search_tool
}Multi-Hop Reasoning with ReAct
ReAct shines on multi-hop questions that require chaining multiple pieces of information. For example: 'Who is the CEO of the company that makes GPT-4, and what year was that company founded?' The agent first searches for the CEO, then uses that result to look up the company's founding year — two separate tool calls chained by reasoning.
ReAct vs. Pure Chain-of-Thought
Chain-of-Thought (CoT) improves reasoning by having the model think step by step, but it cannot access external information. ReAct adds actions to CoT, letting the model fetch real data, run computations, and verify facts. The trade-off: ReAct is slower (multiple API calls) but far more accurate on tasks requiring current or specialized knowledge.
Debugging ReAct with Traces
When a ReAct agent produces the wrong answer, inspect the full Thought/Action/Observation trace. Common failure patterns include: the model inventing an observation instead of calling the tool, parsing errors that skip tool execution, and incorrect reasoning in the Thought step that leads to a wrong action choice.
- Log every message appended to the context
- Check that tool output was actually appended before the next model call
- Verify the action regex matches the model's output format
Quick Check
Test your understanding of the ReAct framework.
Lesson Recap
In this lesson you learned: ReAct interleaves Thought, Action, and Observation in a loop, your application code acts as the runtime that executes tools and appends observations, and max_iterations guards prevent infinite agent loops. Next up we learn to define custom tools so your agent can do exactly what your application needs.
常见问题解答
「ReAct 框架:思考、行动、观察」课时是免费的吗?
是的 — 「ReAct 框架:思考、行动、观察」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 AI Engineering Academy 课程的其余内容,请升级到 CoddyKit PRO。 AI Engineering Academy 课程共包含 4 节课。
「ReAct 框架:思考、行动、观察」这节课中我会学到什么?
了解 ReAct 循环:模型生成 Thought,选择 Action,接收 Observation,并不断重复,直到得出 Final Answer。 你通过在浏览器中直接运行的动手代码来练习 AI Engineering Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 AI Engineering Academy 需要有经验吗?
无需任何先前经验。CoddyKit 上的 AI Engineering Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。
「ReAct 框架:思考、行动、观察」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 AI Engineering Academy 课中编写并运行代码吗?
能。每节 AI Engineering Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- ReAct 框架:思考、行动、观察
- 为您的 Agent 定义工具
- 使用 LangChain 构建 ReAct Agent
- 处理智能体故障与循环