0Pricing
Claude Architect · 课时

结构化错误上下文

失败类型、尝试过的查询、部分结果和替代方案。

结构化错误上下文 是 CoddyKit 上的免费 Claude Architect 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Claude Architect 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Claude Architect 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Why Errors Need Structure

In a multi-agent system, a subagent will eventually hit a failure: a database is down, a query returns nothing, a permission is denied. How that failure is reported decides whether the coordinator can recover intelligently or just gives up.

A generic status like "Operation failed" blocks recovery — the coordinator has no idea what to do next. A structured error context turns a dead end into a routing decision.

This lesson covers the four pillars of a good error context: failure type, attempted query, partial results, and alternatives.

The Generic-Error Anti-Pattern

Compare two error payloads coming back from a subagent or tool.

The generic version tells the coordinator nothing actionable. It can't decide whether to retry, ask the user for more input, or escalate. Silent suppression is even worse — the workflow continues as if data exists when it doesn't.

The structured version names what failed and why, which is the first step toward an intelligent next move.

# Anti-pattern: opaque, un-actionable
return {"isError": True, "message": "Operation failed"}

# Better: structured, routable
return {
    "isError": True,
    "errorCategory": "transient",
    "isRetryable": True,
    "message": "Connection to orders DB timed out after 5s",
}

Pillar 1 — Failure Type

The first job of an error context is to classify the failure. Structured MCP errors carry an errorCategory field with a small, fixed vocabulary:

  • transient — temporary infrastructure fault (timeout, rate limit). Often retryable.
  • validation — the input was malformed.
  • business — a domain rule blocked the action.
  • permission — access was denied.

The companion isRetryable boolean removes guesswork: the coordinator reads it directly instead of inferring intent from a free-text message.

{
  "isError": true,
  "errorCategory": "transient",
  "isRetryable": true,
  "message": "Rate limit hit on inventory service"
}

Failure vs. Empty Result

One distinction trips up architects constantly: an access failure is not the same as a valid empty result.

  • Failure: the lookup couldn't run — DB unreachable, permission denied. This might be worth a retry.
  • Empty: the lookup ran successfully and found zero matches. Retrying changes nothing — the answer is genuinely "none".

Conflating the two leads to pointless retry loops on empty results, or to treating a real outage as "no data found". Always model them as separate states.

def classify(result):
    if result.connection_error:
        return {"isError": True, "errorCategory": "transient",
                "isRetryable": True}
    if not result.rows:                 # ran fine, found nothing
        return {"isError": False, "empty": True, "matches": 0}
    return {"isError": False, "matches": len(result.rows)}

Pillar 2 — Attempted Query

The coordinator did not run the failing operation itself, so it can't see what was tried. Include the attempted_query verbatim in the error context.

This serves two purposes:

  • It lets the coordinator decide whether to retry with the same query or reformulate it (e.g. broaden a filter that was too narrow).
  • It gives the human reviewer, on escalation, the exact reproduction case instead of a vague "search failed".
return {
    "isError": True,
    "errorCategory": "transient",
    "isRetryable": True,
    "attempted_query": {
        "endpoint": "GET /orders",
        "filters": {"customer_id": "C-4821", "status": "shipped"},
    },
    "message": "Orders service returned 503",
}

Pillar 3 — Partial Results

A failure rarely means zero work got done. A research subagent may have gathered 6 of 10 sources before a provider rate-limited it. Throwing all of that away — or aborting the whole workflow — wastes real progress.

Attach whatever was successfully collected as partial_results. The coordinator can then aggregate what exists, annotate the gap, and decide if the remainder is worth another attempt.

Never silently suppress the failure and present partials as if they were complete.

return {
    "isError": True,
    "errorCategory": "transient",
    "isRetryable": True,
    "attempted_query": "fetch 10 sources on 'EU AI Act timelines'",
    "partial_results": collected_sources,   # 6 of 10 gathered
    "message": "Provider rate-limited after 6 sources",
}

Pillar 4 — Alternatives

The most useful error contexts don't just describe the wall — they point at a door. The alternatives field suggests concrete next moves the coordinator (or a human) can take.

Examples: "retry against the read replica", "broaden the date filter", "ask the user for an order number", "escalate to a human with the partial results attached".

This is what turns a structured error from a report into a routing instruction.

return {
    "isError": True,
    "errorCategory": "business",
    "isRetryable": False,
    "attempted_query": "process_refund(order='O-77', amount=620)",
    "partial_results": {"order_total": 620, "customer_verified": True},
    "alternatives": [
        "Refund exceeds $500 policy cap — escalate to human",
        "Offer store credit within auto-approve limit",
    ],
    "message": "Refund blocked by policy threshold",
}

Recover Locally, Escalate Non-Recoverable

Structured context drives a clear policy. Handle transient faults inside the subagent — retry the timeout, back off the rate limit — so the coordinator never even sees a recoverable blip.

Only when a fault is genuinely non-recoverable (policy cap hit, permission denied, retries exhausted) do you escalate upward — and you escalate with the partial results and alternatives attached, not as a bare "failed".

The goal: don't abort the entire workflow because one branch failed.

for attempt in range(3):          # local recovery for transient faults
    res = run_query()
    if not res.get("isError"):
        return res
    if not res.get("isRetryable"):
        break                     # non-recoverable: stop retrying

# escalate upward WITH context, never a bare failure
return escalate(res)

Enforcing the Shape with a Schema

Free-form error dicts drift over time. Enforce the contract with a JSON Schema via a tool / structured output so the subagent must populate the right fields.

Key rule from structured-output design: mark a field required only if it is always present. errorCategory and message are always there — require them. partial_results and alternatives may be absent — leave them optional, or the model will fabricate them to satisfy the schema.

error_schema = {
    "type": "object",
    "properties": {
        "errorCategory": {"enum": ["transient", "validation",
                                    "business", "permission", "other"]},
        "isRetryable": {"type": "boolean"},
        "attempted_query": {"type": "string"},
        "partial_results": {"type": "array"},
        "alternatives": {"type": "array", "items": {"type": "string"}},
        "message": {"type": "string"},
    },
    "required": ["errorCategory", "isRetryable", "message"],
}

Context Doesn't Cross Agent Boundaries for Free

Subagents do not inherit the coordinator's conversation history. So when a subagent fails, the coordinator only knows what the error payload explicitly carries.

That is exactly why attempted_query and partial_results must be in the structured error — there's no shared memory the coordinator can fall back on. The error context is the entire bridge between the two.

Trim it to the relevant fields, but never strip out the four pillars.

# Coordinator delegates; subagent returns ONLY its payload.
# No shared history -> the error context must be self-contained.
results = await asyncio.gather(
    research_subagent("EU AI Act"),
    research_subagent("US AI policy"),
)
for r in results:
    if r.get("isError") and not r["isRetryable"]:
        annotate_coverage_gap(r["attempted_query"], r["partial_results"])

Putting It Together

A production-grade error context lets the coordinator act without re-running anything itself:

  • failure type + isRetryable → retry vs. escalate decision
  • attempted query → reproduce or reformulate
  • partial results → salvage progress, annotate the gap
  • alternatives → the concrete next move

This is the difference between a brittle pipeline that dies on the first hiccup and a resilient system that degrades gracefully and routes around damage.

{
  "isError": true,
  "errorCategory": "permission",
  "isRetryable": false,
  "attempted_query": "SELECT * FROM payroll WHERE dept='ENG'",
  "partial_results": [],
  "alternatives": [
    "Request read grant on payroll schema",
    "Escalate to data-owner for approval"
  ],
  "message": "Access denied to payroll table"
}

Quick Check

A research subagent was asked to gather 10 sources. After collecting 6, the news API rate-limited it (HTTP 429). What should the subagent return to the coordinator?

Recap — Structured Error Context

Key takeaways:

  • Generic errors block recovery; structured errors enable intelligent routing.
  • Always carry the four pillars: failure type, attempted query, partial results, alternatives.
  • Use errorCategory (transient / validation / business / permission) + isRetryable to drive the retry-vs-escalate decision.
  • Distinguish an access failure (maybe retry) from a valid empty result (no matches — retrying won't help).
  • Recover transient faults locally in the subagent; escalate non-recoverable ones with partials attached.
  • Never silently suppress errors and never abort the whole workflow on a single failure.
  • Enforce the shape with a schema, but only require fields that are always present — optional pillars must stay optional to avoid fabrication.

常见问题解答

「结构化错误上下文」课时是免费的吗?

是的 — 「结构化错误上下文」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Claude Architect 课程的其余内容,请升级到 CoddyKit PRO。 Claude Architect 课程共包含 4 节课。

「结构化错误上下文」这节课中我会学到什么?

失败类型、尝试过的查询、部分结果和替代方案。 你通过在浏览器中直接运行的动手代码来练习 Claude Architect,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Claude Architect 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Claude Architect 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。

「结构化错误上下文」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Claude Architect 课中编写并运行代码吗?

能。每节 Claude Architect 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 明确的升级触发条件
  2. 反模式:情绪与置信度评分
  3. 结构化错误上下文
  4. 本地恢复与升级
← 返回 Claude Architect