isError 标志
在 MCP 响应中清晰地表示失败。
isError 标志 是 CoddyKit 上的免费 Claude Architect 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Claude Architect 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Claude Architect 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Why Signalling Failure Matters
When an MCP tool runs, two things can happen: it succeeds, or it fails. The model needs to know which — clearly and unambiguously — to decide what to do next.
If a failure looks like a normal result, the agent may treat garbage as truth, hallucinate a recovery, or silently move on. The fix is a dedicated failure signal: the isError flag.
In this lesson you'll learn how to signal failure cleanly so the agentic loop can route intelligently instead of guessing.
What isError Actually Does
A tool result carries an isError boolean. When isError is true, you're telling Claude: this tool did not produce a valid result — treat the content as a failure report, not data.
This is structurally separate from your tool's normal output. The model can branch on it without parsing prose: success path vs. failure path. That separation is the whole point.
tool_result = {
"type": "tool_result",
"tool_use_id": tool_use.id,
"is_error": True,
"content": "..." # structured failure report
}Generic Errors Block Recovery
The classic anti-pattern is a generic error string like "Operation failed". It tells the model that something went wrong but nothing it can act on.
Can it retry? Was the input malformed? Did the user lack permission? Is there partial data to salvage? A generic message answers none of these — so the agent stalls or improvises badly.
Generic errors block recovery. Structured errors enable intelligent routing.
The Anatomy of a Structured Error
A well-formed MCP error pairs isError: true with a structured body. The exam-standard fields are:
errorCategory— one of transient, validation, business, permissionisRetryable— can the same call succeed if tried again?message— human-readable explanationattempted_query— exactly what the tool tried to dopartial_results— anything usable it managed to gather
Together these let the model decide: retry, reformulate, escalate, or proceed with partial data.
{
"isError": true,
"errorCategory": "transient",
"isRetryable": true,
"message": "Upstream inventory service timed out after 5s",
"attempted_query": "GET /inventory?sku=ABX-19",
"partial_results": null
}errorCategory Drives the Decision
The four categories aren't decoration — each implies a different next action:
- transient — temporary fault (timeout, rate limit). Usually retryable; recover locally.
- validation — bad input. Don't blind-retry; fix the arguments first.
- business — a rule was violated (e.g. refund exceeds policy). Often needs escalation, not retry.
- permission — caller lacks access. Retrying won't help; escalate or request credentials.
The category turns a vague failure into a routing instruction the model can follow.
isRetryable: Don't Make the Model Guess
Whether a failure is worth retrying is often invisible from the message text alone. Make it explicit with isRetryable.
A timeout (transient) is retryable. A malformed argument (validation) is not — retrying the same bad input just fails again. A permission denial is not retryable without new credentials.
By stating isRetryable directly, you keep retry decisions deterministic instead of leaving them to probabilistic text-reading.
{
"isError": true,
"errorCategory": "validation",
"isRetryable": false,
"message": "sku must match pattern ^[A-Z]{3}-[0-9]{2}$; got 'abx19'",
"attempted_query": "lookup_inventory(sku='abx19')"
}attempted_query Preserves Context
When the model decides how to recover, it needs to know what was actually tried. Including attempted_query means the agent can reformulate intelligently instead of repeating the same failing call.
This is part of good error propagation: structured context = failure type, attempted query, partial results, and alternatives. The richer the context, the better the recovery routing.
partial_results: Don't Throw Away Good Data
A tool can fail and still have gathered something useful. A multi-source lookup might return 3 of 5 records before the 4th source times out.
Returning partial_results alongside the error lets the agent proceed with what it has, annotate the gap, and avoid restarting from zero. Discarding partial data on any failure wastes work and degrades answers.
{
"isError": true,
"errorCategory": "transient",
"isRetryable": true,
"message": "3 of 5 sources responded; 2 timed out",
"attempted_query": "search_catalog(term='thermostat')",
"partial_results": [{"id": 11}, {"id": 12}, {"id": 19}]
}Failure vs. a Valid Empty Result
A critical distinction: an access failure is not the same as a valid empty result.
isError: true— the tool couldn't complete (timeout, denied, bad input). Maybe retry or escalate.isError: falsewith empty content — the tool ran fine and the honest answer is "no matches."
Conflating these is a common bug: an empty search marked as an error triggers pointless retries, while a real failure marked as empty hides the problem. Keep them distinct.
{
"isError": false,
"errorCategory": null,
"message": "Query succeeded; 0 orders match customer C-7781",
"results": []
}Recover Locally, Escalate When You Must
The structured signal drives where recovery happens. In a multi-agent system, a subagent should recover transient faults locally — retry the timeout, re-issue the call — and only bubble up failures it truly can't resolve.
When it does escalate, it passes the structured error with partial results so the coordinator can decide: route elsewhere, ask for more identifiers, or surface the gap. Never silently suppress a failure, and never abort the whole workflow over one recoverable fault.
Putting It Together in a Tool
Inside an MCP tool handler, wrap the work and return a structured error on failure instead of letting an exception leak as a generic string.
Notice how each branch sets isError, a category, and a retry hint — giving the agentic loop everything it needs to route the next step deterministically.
def lookup_order(order_id: str):
try:
order = db.fetch(order_id)
if order is None:
return {"isError": False, "results": []} # valid empty
return {"isError": False, "results": [order]}
except TimeoutError as e:
return {
"isError": True,
"errorCategory": "transient",
"isRetryable": True,
"message": str(e),
"attempted_query": f"fetch(order_id={order_id})",
"partial_results": None,
}Quick Check: Choosing the Right Error Shape
A subagent's MCP tool queries a customer's order history. One of three backend shards is unreachable; the other two return 8 orders. What should the tool return?
Recap: Signalling Failure Cleanly
Key takeaways for the isError flag:
- isError: true is the structural signal that a tool did not produce valid data — separate from normal output.
- Pair it with errorCategory (transient / validation / business / permission), isRetryable, message, attempted_query, and partial_results.
- Generic errors like "Operation failed" block recovery; structured errors enable intelligent routing.
- Distinguish an access failure from a valid empty result — never mark "no matches" as an error.
- Recover transient faults locally; escalate non-recoverable ones with partial results. Avoid silent suppression and avoid aborting the whole workflow on one fault.
常见问题解答
「isError 标志」课时是免费的吗?
是的 — 「isError 标志」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Claude Architect 课程的其余内容,请升级到 CoddyKit PRO。 Claude Architect 课程共包含 4 节课。
「isError 标志」这节课中我会学到什么?
在 MCP 响应中清晰地表示失败。 你通过在浏览器中直接运行的动手代码来练习 Claude Architect,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Claude Architect 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Claude Architect 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。
「isError 标志」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Claude Architect 课中编写并运行代码吗?
能。每节 Claude Architect 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- isError 标志
- 错误类别
- 可重试元数据与部分结果
- 反模式:通用错误消息