The isError Flag
Signalling failure cleanly in MCP responses.
The isError Flag is a free Claude Architect lesson on CoddyKit — lesson 1 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Claude Architect learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
Why Signalling Failure Matters
When an MCP tool runs, two things can happen: it succeeds, or it fails. The model needs to know which — clearly and unambiguously — to decide what to do next.
If a failure looks like a normal result, the agent may treat garbage as truth, hallucinate a recovery, or silently move on. The fix is a dedicated failure signal: the isError flag.
In this lesson you'll learn how to signal failure cleanly so the agentic loop can route intelligently instead of guessing.
What isError Actually Does
A tool result carries an isError boolean. When isError is true, you're telling Claude: this tool did not produce a valid result — treat the content as a failure report, not data.
This is structurally separate from your tool's normal output. The model can branch on it without parsing prose: success path vs. failure path. That separation is the whole point.
tool_result = {
"type": "tool_result",
"tool_use_id": tool_use.id,
"is_error": True,
"content": "..." # structured failure report
}Generic Errors Block Recovery
The classic anti-pattern is a generic error string like "Operation failed". It tells the model that something went wrong but nothing it can act on.
Can it retry? Was the input malformed? Did the user lack permission? Is there partial data to salvage? A generic message answers none of these — so the agent stalls or improvises badly.
Generic errors block recovery. Structured errors enable intelligent routing.
The Anatomy of a Structured Error
A well-formed MCP error pairs isError: true with a structured body. The exam-standard fields are:
errorCategory— one of transient, validation, business, permissionisRetryable— can the same call succeed if tried again?message— human-readable explanationattempted_query— exactly what the tool tried to dopartial_results— anything usable it managed to gather
Together these let the model decide: retry, reformulate, escalate, or proceed with partial data.
{
"isError": true,
"errorCategory": "transient",
"isRetryable": true,
"message": "Upstream inventory service timed out after 5s",
"attempted_query": "GET /inventory?sku=ABX-19",
"partial_results": null
}errorCategory Drives the Decision
The four categories aren't decoration — each implies a different next action:
- transient — temporary fault (timeout, rate limit). Usually retryable; recover locally.
- validation — bad input. Don't blind-retry; fix the arguments first.
- business — a rule was violated (e.g. refund exceeds policy). Often needs escalation, not retry.
- permission — caller lacks access. Retrying won't help; escalate or request credentials.
The category turns a vague failure into a routing instruction the model can follow.
isRetryable: Don't Make the Model Guess
Whether a failure is worth retrying is often invisible from the message text alone. Make it explicit with isRetryable.
A timeout (transient) is retryable. A malformed argument (validation) is not — retrying the same bad input just fails again. A permission denial is not retryable without new credentials.
By stating isRetryable directly, you keep retry decisions deterministic instead of leaving them to probabilistic text-reading.
{
"isError": true,
"errorCategory": "validation",
"isRetryable": false,
"message": "sku must match pattern ^[A-Z]{3}-[0-9]{2}$; got 'abx19'",
"attempted_query": "lookup_inventory(sku='abx19')"
}attempted_query Preserves Context
When the model decides how to recover, it needs to know what was actually tried. Including attempted_query means the agent can reformulate intelligently instead of repeating the same failing call.
This is part of good error propagation: structured context = failure type, attempted query, partial results, and alternatives. The richer the context, the better the recovery routing.
partial_results: Don't Throw Away Good Data
A tool can fail and still have gathered something useful. A multi-source lookup might return 3 of 5 records before the 4th source times out.
Returning partial_results alongside the error lets the agent proceed with what it has, annotate the gap, and avoid restarting from zero. Discarding partial data on any failure wastes work and degrades answers.
{
"isError": true,
"errorCategory": "transient",
"isRetryable": true,
"message": "3 of 5 sources responded; 2 timed out",
"attempted_query": "search_catalog(term='thermostat')",
"partial_results": [{"id": 11}, {"id": 12}, {"id": 19}]
}Failure vs. a Valid Empty Result
A critical distinction: an access failure is not the same as a valid empty result.
isError: true— the tool couldn't complete (timeout, denied, bad input). Maybe retry or escalate.isError: falsewith empty content — the tool ran fine and the honest answer is "no matches."
Conflating these is a common bug: an empty search marked as an error triggers pointless retries, while a real failure marked as empty hides the problem. Keep them distinct.
{
"isError": false,
"errorCategory": null,
"message": "Query succeeded; 0 orders match customer C-7781",
"results": []
}Recover Locally, Escalate When You Must
The structured signal drives where recovery happens. In a multi-agent system, a subagent should recover transient faults locally — retry the timeout, re-issue the call — and only bubble up failures it truly can't resolve.
When it does escalate, it passes the structured error with partial results so the coordinator can decide: route elsewhere, ask for more identifiers, or surface the gap. Never silently suppress a failure, and never abort the whole workflow over one recoverable fault.
Putting It Together in a Tool
Inside an MCP tool handler, wrap the work and return a structured error on failure instead of letting an exception leak as a generic string.
Notice how each branch sets isError, a category, and a retry hint — giving the agentic loop everything it needs to route the next step deterministically.
def lookup_order(order_id: str):
try:
order = db.fetch(order_id)
if order is None:
return {"isError": False, "results": []} # valid empty
return {"isError": False, "results": [order]}
except TimeoutError as e:
return {
"isError": True,
"errorCategory": "transient",
"isRetryable": True,
"message": str(e),
"attempted_query": f"fetch(order_id={order_id})",
"partial_results": None,
}Quick Check: Choosing the Right Error Shape
A subagent's MCP tool queries a customer's order history. One of three backend shards is unreachable; the other two return 8 orders. What should the tool return?
Recap: Signalling Failure Cleanly
Key takeaways for the isError flag:
- isError: true is the structural signal that a tool did not produce valid data — separate from normal output.
- Pair it with errorCategory (transient / validation / business / permission), isRetryable, message, attempted_query, and partial_results.
- Generic errors like "Operation failed" block recovery; structured errors enable intelligent routing.
- Distinguish an access failure from a valid empty result — never mark "no matches" as an error.
- Recover transient faults locally; escalate non-recoverable ones with partial results. Avoid silent suppression and avoid aborting the whole workflow on one fault.
Frequently asked questions
Is the “The isError Flag” lesson free?
Yes — the full text of “The isError Flag” is free to read here on the web, and the Claude Architect course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Claude Architect course, upgrade to CoddyKit PRO.
What will I learn in “The isError Flag”?
Signalling failure cleanly in MCP responses. You practise Claude Architect with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start Claude Architect?
No prior experience is required. Claude Architect on CoddyKit is structured for beginners through advanced learners; this is — lesson 1 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “The isError Flag” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this Claude Architect lesson?
Yes. Every Claude Architect lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.