停止原因详解
end_turn、tool_use、max_tokens 和 stop_sequence。
停止原因详解 是 CoddyKit 上的免费 Claude Architect 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Claude Architect 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Claude Architect 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Why Stop Reasons Matter
Every time you call the Claude API, the response comes back with a stop_reason field. It tells you why the model stopped generating.
This one field drives your whole control flow. A reliable agent inspects stop_reason after each turn and decides what to do next based on it.
There are four values you must know: end_turn, tool_use, max_tokens, and stop_sequence. Let's learn each one.
Where to Find It
The stop_reason lives on the response object returned by messages.create.
Read it directly. Do not scan the text output for words like "done" or "finished" to decide what happened. The model controls stop_reason; text is just content.
from anthropic import Anthropic
client = Anthropic()
response = client.messages.create(
model="claude-opus-4-8",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello"}],
)
print(response.stop_reason) # "end_turn"end_turn — Complete
end_turn means Claude finished its response naturally. It said everything it wanted to say.
This is the normal completion signal. In an agent loop, end_turn is your cue to stop looping and return the answer to the user.
if response.stop_reason == "end_turn":
# Claude is done. Return the answer.
print(response.content[0].text)tool_use — Run a Tool
tool_use means Claude wants to call one of the tools you gave it. The response is not finished — Claude is waiting for a tool result.
Your job: run the requested tool, append the result to the message history, and call the API again so Claude can continue.
if response.stop_reason == "tool_use":
tool_call = next(b for b in response.content if b.type == "tool_use")
result = run_tool(tool_call.name, tool_call.input)
# Append the result and loop again (next scenes show how)max_tokens — Truncated
max_tokens means the response was cut off because it hit the max_tokens limit you set in the request.
The output is incomplete — it stopped mid-thought, not because Claude was done. The fix is to raise max_tokens, or stream the response for very long outputs.
Never treat max_tokens as a successful completion.
if response.stop_reason == "max_tokens":
# Output was truncated. Retry with a higher max_tokens,
# or use client.messages.stream(...) for long outputs.
print("Response was cut off — increase max_tokens.")stop_sequence — Custom Stop
stop_sequence means Claude hit a custom stop string that you defined in your request via stop_sequences.
Generation halts the moment that string is produced. This is useful when you want the model to stop at a known boundary, such as "###" or "END".
response = client.messages.create(
model="claude-opus-4-8",
max_tokens=1024,
stop_sequences=["###"],
messages=[{"role": "user", "content": "List three colors, then ###"}],
)
if response.stop_reason == "stop_sequence":
print("Stopped at a custom sequence.")The Four at a Glance
Here is the full set for this lesson:
- end_turn — Claude finished naturally. Stop the loop.
- tool_use — Claude wants a tool. Run it, append the result, call again.
- max_tokens — Output truncated. Raise the limit or stream.
- stop_sequence — Hit a custom stop string you defined.
Two of these (end_turn, stop_sequence) mean the turn is complete. One (tool_use) means continue. One (max_tokens) means something went wrong with your limit.
The Agentic Loop
An agent is just a loop driven by stop_reason:
Send a request, inspect stop_reason. If it is tool_use, run the tools, append the results to the history, and repeat. Keep going until stop_reason is end_turn.
The model keeps no state between calls, so you must send the full message history every turn.
messages = [{"role": "user", "content": user_input}]
while True:
response = client.messages.create(
model="claude-opus-4-8",
max_tokens=4096,
tools=tools,
messages=messages,
)
if response.stop_reason == "end_turn":
break
if response.stop_reason == "tool_use":
messages.append({"role": "assistant", "content": response.content})
results = run_tools(response.content)
messages.append({"role": "user", "content": results})Terminate on the Signal, Not the Text
The single most important rule: terminate the loop on stop_reason, never by parsing text for words like "done", "finished", or "complete".
Text-matching is fragile — the model might say "I'm done thinking, let me check one more file" and your code would stop too early. The stop_reason is the model's structured, reliable signal.
Iteration Caps Are a Safety Net
You may add a maximum iteration count to your loop so a runaway agent cannot loop forever. That is good practice.
But an iteration cap is a safety net, never the primary way you stop. The primary stop is always end_turn. Decisions about when work is done are model-driven; the cap only catches the rare case where something goes wrong.
MAX_ITERS = 25 # safety net only
for _ in range(MAX_ITERS):
response = client.messages.create(...)
if response.stop_reason == "end_turn":
break # the real, primary stop
# ... handle tool_use ...
else:
log.warning("Hit iteration cap — investigate.")Putting It Together
A robust handler branches on every stop reason explicitly:
end_turn→ return the result.tool_use→ execute tools, append results, continue.max_tokens→ the output is truncated; raise the limit and retry rather than using a partial answer.stop_sequence→ handle the known boundary you defined.
Handling all branches is what separates a reliable agent from one that silently breaks on edge cases.
def handle(response):
sr = response.stop_reason
if sr == "end_turn":
return finish(response)
if sr == "tool_use":
return continue_with_tools(response)
if sr == "max_tokens":
return retry_with_more_tokens(response)
if sr == "stop_sequence":
return handle_boundary(response)Quick Check
An agent loop returns a response with stop_reason == "tool_use". What is the correct next action?
Recap
Key takeaways:
- end_turn — complete; stop the loop and return.
- tool_use — run the tool, append the result, call again.
- max_tokens — truncated; raise the limit or stream, don't use the partial output.
- stop_sequence — hit a custom stop string you defined.
Always drive your control flow from stop_reason, never from parsing text. Terminate on end_turn; keep iteration caps as a safety net only. Master this and the agentic loop becomes simple and reliable.
用 AI 导师学习 Python — 免费
在浏览器中编写并运行真实代码,获得全天候 AI 导师的即时帮助,并在网页或应用中继续学习。
- 课程
- 26
- 课程
- 104
常见问题解答
「停止原因详解」课时是免费的吗?
是的 — 「停止原因详解」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Claude Architect 课程的其余内容,请升级到 CoddyKit PRO。 Claude Architect 课程共包含 4 节课。
「停止原因详解」这节课中我会学到什么?
end_turn、tool_use、max_tokens 和 stop_sequence。 你通过在浏览器中直接运行的动手代码来练习 Claude Architect,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Claude Architect 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Claude Architect 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。
「停止原因详解」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Claude Architect 课中编写并运行代码吗?
能。每节 Claude Architect 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。