0Pricing
Claude Architect · 课时

错误类别

瞬态错误、验证错误、业务错误和权限错误。

错误类别 是 CoddyKit 上的免费 Claude Architect 课时。 这是第 2 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Claude Architect 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Claude Architect 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Why Error Categories Matter

When a tool fails inside an agent, how it reports the failure decides whether Claude can recover. A generic status like "Operation failed" tells the model nothing — it can't tell a temporary blip from a permanent block, so it either retries blindly or gives up.

The fix is a structured error: a small object that names the kind of failure. The core field is errorCategory, which takes one of four values:

  • transient — temporary, likely to succeed on retry
  • validation — the request was malformed
  • business — a domain rule blocked it
  • permission — access was denied

This lesson teaches you to recognize each category and route on it.

The Shape of a Structured Error

A well-designed MCP tool error carries more than a message. It flags itself as an error, names a category, and says whether retrying could help.

Notice the fields: isError marks it as a failure, errorCategory classifies it, and isRetryable gives an explicit yes/no on retrying. The extra fields — attempted_query and partial_results — let the coordinator recover intelligently instead of starting over.

# A structured tool error returned to Claude
{
  "isError": True,
  "errorCategory": "transient",   # transient | validation | business | permission
  "isRetryable": True,
  "message": "Upstream inventory service timed out after 5s",
  "attempted_query": "SELECT stock FROM inventory WHERE sku='A-19'",
  "partial_results": []
}

Transient Errors

Transient errors are temporary and not your request's fault: a timeout, a brief network drop, an upstream service that is momentarily overloaded. The same call, made again a moment later, is likely to succeed.

These map to isRetryable: true. The right move is almost always to retry locally inside the subagent — recover the fault where it happened, with a short backoff, rather than aborting the whole workflow over a hiccup.

{
  "isError": True,
  "errorCategory": "transient",
  "isRetryable": True,
  "message": "503 from payment gateway; service temporarily overloaded"
}

Validation Errors

Validation errors mean the request itself was malformed: a missing required field, a wrong type, an out-of-range value, a bad date format. The tool never got far enough to do real work.

Blindly retrying the identical request will fail again — so isRetryable is usually false. But the error is still recoverable: Claude can read the message, fix the input, and call again. The richer the message (which field, what was expected), the faster the model corrects itself.

{
  "isError": True,
  "errorCategory": "validation",
  "isRetryable": False,
  "message": "Field 'order_date' must be ISO-8601 (got '13/2026'); expected e.g. 2026-02-13"
}

Business Errors

Business errors happen when the request is well-formed and the caller is allowed, but a domain rule forbids the action. Think: refund exceeds the policy ceiling, account balance too low, item out of stock, booking past the cutoff time.

This is not a bug to retry — it's a real-world constraint. isRetryable is false. The agent should surface the rule clearly, and if the workflow can't proceed within policy, escalate rather than try to force the action.

{
  "isError": True,
  "errorCategory": "business",
  "isRetryable": False,
  "message": "Refund of $640 exceeds the $500 auto-approval limit; requires human approval"
}

Permission Errors

Permission errors mean access was denied: the credential, token, or role does not authorize this action or resource. The request may be perfectly valid in shape and intent — it's simply not allowed.

Retrying the same call with the same credential will keep failing, so isRetryable is false. The agent should not loop; it should report the access gap and, when appropriate, escalate to a human or a privileged path.

{
  "isError": True,
  "errorCategory": "permission",
  "isRetryable": False,
  "message": "API key lacks scope 'orders:write'; cannot process refund"
}

Routing on the Category

The whole point of categorizing is to route. Once the coordinator reads errorCategory, the recovery path is deterministic — it doesn't have to guess by reading prose.

A clean mapping:

  • transient → retry locally with backoff
  • validation → fix the input, then retry
  • business → respect the rule; escalate if blocked
  • permission → stop; report / escalate the access gap

This is intelligent routing that a generic "failed" status simply cannot support.

def route(err):
    cat = err["errorCategory"]
    if cat == "transient":
        return retry_with_backoff(err)      # recover in the subagent
    if cat == "validation":
        return repair_input_and_retry(err)  # not the same request twice
    if cat == "business":
        return escalate_if_blocked(err)     # respect the policy
    if cat == "permission":
        return report_access_gap(err)       # stop; do not loop

isRetryable Is the Fast Path

You don't always have to branch on all four categories. The isRetryable flag is a quick gate: true means "trying again could plausibly work," false means "the same call will fail the same way."

As a rule of thumb, only transient errors are retryable as-is. Validation needs the input changed first; business and permission won't change on retry at all. Treat isRetryable: false as a hard signal to stop retrying and choose a different action — repair, escalate, or report.

if err["isError"] and err["isRetryable"]:
    # transient: safe to try again with backoff
    result = retry_with_backoff(err)
else:
    # validation / business / permission: retrying won't help
    result = handle_non_retryable(err)

Access Failure vs Empty Result

A subtle but exam-critical distinction: an access failure is not the same as a valid empty result.

If a lookup tool can't reach the database, that's a transient error — maybe retry. If it reaches the database fine and there are simply no matching rows, that is a successful call returning zero results — isError is false. Conflating the two leads to pointless retries on "no matches found," or worse, silently treating a real failure as "nothing there."

# NOT an error — a valid empty result
{ "isError": False, "results": [], "message": "No orders match customer 8842" }

# An error — could not even run the query
{ "isError": True, "errorCategory": "transient",
  "isRetryable": True, "message": "Connection refused to orders DB" }

Carry Partial Results and Context

When a non-recoverable failure must propagate up, send structured context with it — not just "it broke." Include the failure type, the attempted_query, any partial_results already gathered, and viable alternatives.

In a multi-agent research system, this is what lets the coordinator annotate coverage gaps and still assemble a useful answer from what succeeded. The two failure modes to avoid: silent suppression (swallowing the error) and aborting the entire workflow because one of ten subtasks failed.

{
  "isError": True,
  "errorCategory": "permission",
  "isRetryable": False,
  "message": "No access to EU sales shard",
  "attempted_query": "sales WHERE region='EU' AND year=2026",
  "partial_results": [{"region": "US", "total": 4200000}]
}

Recover Local, Escalate Non-Recoverable

Put the two halves together into one operating principle for agent error handling:

  • Recover transient faults locally — retry inside the subagent with backoff; don't bubble a timeout all the way to the human.
  • Escalate non-recoverable failures with partial results — when business rules, permissions, or absent data block progress, hand up the structured context so a human or coordinator can decide.

Escalation triggers should be concrete (policy gaps, threshold violations, no progress after attempts, explicit human requests) — never based on sentiment or a model's self-rated confidence score.

Quick Check: Choosing the Recovery Path

Test your routing instinct on a realistic agent failure.

Recap: Four Categories, One Discipline

Structured errors turn a dead end into a decision. Remember the four categories and their default routes:

  • transient (isRetryable: true) → retry locally with backoff
  • validation → fix the input, then retry
  • business → respect the rule; escalate if blocked
  • permission → stop; report or escalate the access gap

Always distinguish an access failure from a valid empty result. When propagating a failure, carry attempted_query and partial_results — never suppress silently and never abort the whole workflow on one failure. Generic errors block recovery; structured errors with a category and isRetryable enable it.

常见问题解答

「错误类别」课时是免费的吗?

是的 — 「错误类别」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Claude Architect 课程的其余内容,请升级到 CoddyKit PRO。 Claude Architect 课程共包含 4 节课。

「错误类别」这节课中我会学到什么?

瞬态错误、验证错误、业务错误和权限错误。 你通过在浏览器中直接运行的动手代码来练习 Claude Architect,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Claude Architect 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Claude Architect 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 2 节课,共 4 节。

「错误类别」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Claude Architect 课中编写并运行代码吗?

能。每节 Claude Architect 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. isError 标志
  2. 错误类别
  3. 可重试元数据与部分结果
  4. 反模式:通用错误消息
← 返回 Claude Architect