0Pricing
Claude Architect · 강의

오류 범주

일시적 오류, 검증 오류, 업무 오류 및 권한 오류를 다룹니다

오류 범주은(는) CoddyKit의 무료 Claude Architect 강의입니다. 이것은 4개 중 2번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 Claude Architect 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. Claude Architect 강의에는 총 4개의 강의가 포함되어 있습니다.

이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.

Why Error Categories Matter

When a tool fails inside an agent, how it reports the failure decides whether Claude can recover. A generic status like "Operation failed" tells the model nothing — it can't tell a temporary blip from a permanent block, so it either retries blindly or gives up.

The fix is a structured error: a small object that names the kind of failure. The core field is errorCategory, which takes one of four values:

  • transient — temporary, likely to succeed on retry
  • validation — the request was malformed
  • business — a domain rule blocked it
  • permission — access was denied

This lesson teaches you to recognize each category and route on it.

The Shape of a Structured Error

A well-designed MCP tool error carries more than a message. It flags itself as an error, names a category, and says whether retrying could help.

Notice the fields: isError marks it as a failure, errorCategory classifies it, and isRetryable gives an explicit yes/no on retrying. The extra fields — attempted_query and partial_results — let the coordinator recover intelligently instead of starting over.

# A structured tool error returned to Claude
{
  "isError": True,
  "errorCategory": "transient",   # transient | validation | business | permission
  "isRetryable": True,
  "message": "Upstream inventory service timed out after 5s",
  "attempted_query": "SELECT stock FROM inventory WHERE sku='A-19'",
  "partial_results": []
}

Transient Errors

Transient errors are temporary and not your request's fault: a timeout, a brief network drop, an upstream service that is momentarily overloaded. The same call, made again a moment later, is likely to succeed.

These map to isRetryable: true. The right move is almost always to retry locally inside the subagent — recover the fault where it happened, with a short backoff, rather than aborting the whole workflow over a hiccup.

{
  "isError": True,
  "errorCategory": "transient",
  "isRetryable": True,
  "message": "503 from payment gateway; service temporarily overloaded"
}

Validation Errors

Validation errors mean the request itself was malformed: a missing required field, a wrong type, an out-of-range value, a bad date format. The tool never got far enough to do real work.

Blindly retrying the identical request will fail again — so isRetryable is usually false. But the error is still recoverable: Claude can read the message, fix the input, and call again. The richer the message (which field, what was expected), the faster the model corrects itself.

{
  "isError": True,
  "errorCategory": "validation",
  "isRetryable": False,
  "message": "Field 'order_date' must be ISO-8601 (got '13/2026'); expected e.g. 2026-02-13"
}

Business Errors

Business errors happen when the request is well-formed and the caller is allowed, but a domain rule forbids the action. Think: refund exceeds the policy ceiling, account balance too low, item out of stock, booking past the cutoff time.

This is not a bug to retry — it's a real-world constraint. isRetryable is false. The agent should surface the rule clearly, and if the workflow can't proceed within policy, escalate rather than try to force the action.

{
  "isError": True,
  "errorCategory": "business",
  "isRetryable": False,
  "message": "Refund of $640 exceeds the $500 auto-approval limit; requires human approval"
}

Permission Errors

Permission errors mean access was denied: the credential, token, or role does not authorize this action or resource. The request may be perfectly valid in shape and intent — it's simply not allowed.

Retrying the same call with the same credential will keep failing, so isRetryable is false. The agent should not loop; it should report the access gap and, when appropriate, escalate to a human or a privileged path.

{
  "isError": True,
  "errorCategory": "permission",
  "isRetryable": False,
  "message": "API key lacks scope 'orders:write'; cannot process refund"
}

Routing on the Category

The whole point of categorizing is to route. Once the coordinator reads errorCategory, the recovery path is deterministic — it doesn't have to guess by reading prose.

A clean mapping:

  • transient → retry locally with backoff
  • validation → fix the input, then retry
  • business → respect the rule; escalate if blocked
  • permission → stop; report / escalate the access gap

This is intelligent routing that a generic "failed" status simply cannot support.

def route(err):
    cat = err["errorCategory"]
    if cat == "transient":
        return retry_with_backoff(err)      # recover in the subagent
    if cat == "validation":
        return repair_input_and_retry(err)  # not the same request twice
    if cat == "business":
        return escalate_if_blocked(err)     # respect the policy
    if cat == "permission":
        return report_access_gap(err)       # stop; do not loop

isRetryable Is the Fast Path

You don't always have to branch on all four categories. The isRetryable flag is a quick gate: true means "trying again could plausibly work," false means "the same call will fail the same way."

As a rule of thumb, only transient errors are retryable as-is. Validation needs the input changed first; business and permission won't change on retry at all. Treat isRetryable: false as a hard signal to stop retrying and choose a different action — repair, escalate, or report.

if err["isError"] and err["isRetryable"]:
    # transient: safe to try again with backoff
    result = retry_with_backoff(err)
else:
    # validation / business / permission: retrying won't help
    result = handle_non_retryable(err)

Access Failure vs Empty Result

A subtle but exam-critical distinction: an access failure is not the same as a valid empty result.

If a lookup tool can't reach the database, that's a transient error — maybe retry. If it reaches the database fine and there are simply no matching rows, that is a successful call returning zero results — isError is false. Conflating the two leads to pointless retries on "no matches found," or worse, silently treating a real failure as "nothing there."

# NOT an error — a valid empty result
{ "isError": False, "results": [], "message": "No orders match customer 8842" }

# An error — could not even run the query
{ "isError": True, "errorCategory": "transient",
  "isRetryable": True, "message": "Connection refused to orders DB" }

Carry Partial Results and Context

When a non-recoverable failure must propagate up, send structured context with it — not just "it broke." Include the failure type, the attempted_query, any partial_results already gathered, and viable alternatives.

In a multi-agent research system, this is what lets the coordinator annotate coverage gaps and still assemble a useful answer from what succeeded. The two failure modes to avoid: silent suppression (swallowing the error) and aborting the entire workflow because one of ten subtasks failed.

{
  "isError": True,
  "errorCategory": "permission",
  "isRetryable": False,
  "message": "No access to EU sales shard",
  "attempted_query": "sales WHERE region='EU' AND year=2026",
  "partial_results": [{"region": "US", "total": 4200000}]
}

Recover Local, Escalate Non-Recoverable

Put the two halves together into one operating principle for agent error handling:

  • Recover transient faults locally — retry inside the subagent with backoff; don't bubble a timeout all the way to the human.
  • Escalate non-recoverable failures with partial results — when business rules, permissions, or absent data block progress, hand up the structured context so a human or coordinator can decide.

Escalation triggers should be concrete (policy gaps, threshold violations, no progress after attempts, explicit human requests) — never based on sentiment or a model's self-rated confidence score.

Quick Check: Choosing the Recovery Path

Test your routing instinct on a realistic agent failure.

Recap: Four Categories, One Discipline

Structured errors turn a dead end into a decision. Remember the four categories and their default routes:

  • transient (isRetryable: true) → retry locally with backoff
  • validation → fix the input, then retry
  • business → respect the rule; escalate if blocked
  • permission → stop; report or escalate the access gap

Always distinguish an access failure from a valid empty result. When propagating a failure, carry attempted_query and partial_results — never suppress silently and never abort the whole workflow on one failure. Generic errors block recovery; structured errors with a category and isRetryable enable it.

자주 묻는 질문

“오류 범주” 강의는 무료인가요?

네 — “오류 범주” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 Claude Architect 강의 전체를 잠금 해제할 수 있습니다. Claude Architect 강의에는 총 4개의 강의가 포함되어 있습니다.

“오류 범주”에서 뭘 배우나요?

일시적 오류, 검증 오류, 업무 오류 및 권한 오류를 다룹니다 브라우저에서 직접 실행하는 실습 코드로 Claude Architect을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.

Claude Architect을(를) 시작하는 데 경험이 필요한가요?

사전 경험은 필요하지 않습니다. CoddyKit의 Claude Architect은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 2번째 강의입니다.

“오류 범주” 강의는 얼마나 걸리나요?

대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.

이 Claude Architect 강의에서 코드를 작성하고 실행할 수 있나요?

네. 모든 Claude Architect 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.

이 강의의 모든 강의

  1. isError 플래그
  2. 오류 범주
  3. 재시도 가능 메타데이터 및 부분 결과
  4. 안티 패턴: 일반적인 오류 메시지
← Claude Architect(으)로 돌아가기