0Pricing
Claude Architect · Lesson

Parallel Subagent Spawning

Multiple Task calls in one response run concurrently.

Parallel Subagent Spawning is a free Claude Architect lesson on CoddyKit — lesson 4 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Claude Architect learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.

Why Parallelism Matters

In a hub-and-spoke multi-agent system, a coordinator decomposes a task, delegates pieces to subagents, then aggregates the results. The key performance lever: multiple Task calls emitted in one response run concurrently.

If a research job needs three independent lookups, you don't have to do them one after another. Spawn all three in a single turn and they execute in parallel, collapsing three round-trips into one wave.

The Parallel Rule

The rule is precise: every Task call the model emits in the same assistant response is dispatched at the same time. There is no special API flag — concurrency is a property of how many Task calls share one response.

  • Three Task calls in one response → three subagents run in parallel.
  • One Task call, wait for the result, then another Task call next turn → sequential.

So parallelism is a decomposition decision, not a configuration toggle.

Coordinator Must Allow Task

A coordinator can only spawn subagents if its allowedTools includes "Task". Without it, the model has no mechanism to delegate and will try to do everything itself.

This is part of least-privilege design: only the coordinator gets Task; leaf subagents usually do not, so they can't recursively spawn more agents.

from claude_agent_sdk import AgentDefinition

coordinator = AgentDefinition(
    name="research_coordinator",
    description="Decomposes a research question and delegates to subagents in parallel.",
    system_prompt="Break the question into independent sub-questions and spawn one subagent per topic IN A SINGLE RESPONSE.",
    allowed_tools=["Task"],  # required to spawn subagents
)

Subagents Don't Inherit History

This is the single most-tested fact about subagents: a subagent does NOT inherit the coordinator's conversation history. It starts with a clean context.

Therefore all context a subagent needs must be passed explicitly in its prompt — the question, relevant constraints, prior findings, output format. If you assume the subagent "already knows" what the coordinator discussed, it will hallucinate or under-perform.

Self-Contained Subagent Prompts

Because context isn't inherited, each parallel Task prompt should be fully self-contained. Notice how each spawned subagent below carries its own topic and the same explicit instructions — nothing is left implicit.

# Coordinator emits these THREE Task calls in ONE response -> they run in parallel
tasks = [
    Task(agent="researcher",
         prompt="Topic: 2024 EV battery cost trends.\n"
                "Return: 3 findings, each with source URL + publication date."),
    Task(agent="researcher",
         prompt="Topic: solid-state battery timelines.\n"
                "Return: 3 findings, each with source URL + publication date."),
    Task(agent="researcher",
         prompt="Topic: lithium supply constraints 2025.\n"
                "Return: 3 findings, each with source URL + publication date."),
]

When Parallel Is Correct

Parallel spawning is the right move when sub-tasks are independent — none needs the output of another.

  • Multi-topic research: each topic is its own lane.
  • Per-file local code review: review each file independently in parallel, then a separate cross-file integration pass.
  • Fan-out lookups: several catalogs or APIs queried at once.

If task B needs A's result, that's a sequential dependency — use a fixed pipeline / prompt chaining instead, not parallel spawning.

Parallel vs. Sequential Decomposition

Match the decomposition pattern to the work:

  • Parallel Task fan-out → independent subtasks, latency-sensitive, results merged at the end.
  • Fixed pipeline / prompt chaining → known sequential steps where each feeds the next.
  • Adaptive decomposition → open-ended investigation where the next step depends on what you just found.

Forcing inherently sequential steps into one parallel wave produces agents working with missing inputs — a correctness bug, not a speedup.

# Sequential: step 2 NEEDS step 1's output -> chain, do NOT parallelize
# 1) fetch the schema, THEN 2) extract rows against that schema
schema = Task(agent="schema_reader", prompt="Return the orders table schema.")
# ...wait for result, then next turn:
extract = Task(agent="extractor",
               prompt=f"Using this schema:\n{schema}\nExtract Q1 orders.")

Aggregating Parallel Results

After a parallel wave returns, the coordinator's job is to aggregate: merge the findings, deduplicate, and resolve conflicts. Don't arbitrarily pick one number when two subagents disagree — annotate the conflict (dates often resolve apparent contradictions).

The coordinator also routes follow-ups: if one lane came back thin, it can spawn a targeted second wave.

# After the parallel wave, the coordinator receives all three result blocks
# and synthesizes. Keep claim -> source mappings for provenance.
synthesis_prompt = (
    "You have results from 3 parallel researchers below.\n"
    "Merge into one report. For each claim keep its source URL + date.\n"
    "If two sources conflict, annotate both rather than picking one."
)

Partial Failure in a Parallel Wave

When you fan out, one lane may fail while others succeed. Do not abort the whole workflow on a single failure, and never silently suppress it.

  • Recover transient faults locally inside the subagent (retry).
  • Escalate non-recoverable failures with partial results and structured context: failure type, attempted query, what did succeed.

The coordinator should annotate coverage gaps ("topic 2 unavailable") rather than presenting incomplete output as complete.

# A subagent returns STRUCTURED context on failure, not a silent drop
{
    "isError": True,
    "errorCategory": "transient",      # transient|validation|business|permission
    "isRetryable": True,
    "message": "Search API timed out",
    "attempted_query": "solid-state battery timelines",
    "partial_results": ["1 finding retrieved before timeout"],
}

Least-Privilege Subagents

Each subagent's allowed_tools should be scoped to its role — least privilege. A researcher needs read/search tools; it does not need Task (no recursive spawning) and it does not need write or refund tools.

Remember the tool-count guideline: 4-5 tools per agent is optimal; 18+ degrades selection reliability. Narrow, role-scoped subagents pick the right tool far more reliably than one overloaded mega-agent.

researcher = AgentDefinition(
    name="researcher",
    description="Investigates ONE topic and returns findings with citations.",
    system_prompt="Research the given topic. Return findings with source URL + date.",
    allowed_tools=["WebSearch", "WebFetch", "Read"],  # no Task: cannot re-spawn
)

The Agentic Loop Still Governs

Parallelism doesn't change the control flow. The coordinator still runs the standard agentic loop: send request → inspect stop_reason → if tool_use, run the tools (the parallel subagents) and append their results to history → repeat until end_turn.

Terminate on stop_reason, never by parsing text for words like "done". Iteration caps are a safety net, not the primary stop mechanism. Whether one Task or three ran, the loop logic is identical.

Quick Check

A coordinator must research three independent topics as fast as possible. Which approach correctly spawns the subagents in parallel?

Recap

Key takeaways for parallel subagent spawning:

  • Multiple Task calls in one response run in parallel — there's no flag, it's a decomposition choice.
  • The coordinator's allowedTools must include "Task".
  • Subagents do not inherit history — pass all context explicitly in each self-contained prompt.
  • Use parallel for independent subtasks; use a pipeline/chaining for sequential dependencies.
  • On partial failure: recover transient faults locally, escalate non-recoverable ones with partial results — never abort the whole workflow or suppress silently.
  • Scope subagents with least privilege (4-5 tools), and keep terminating on stop_reason, never on parsed text.

Frequently asked questions

Is the “Parallel Subagent Spawning” lesson free?

Yes — the full text of “Parallel Subagent Spawning” is free to read here on the web, and the Claude Architect course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Claude Architect course, upgrade to CoddyKit PRO.

What will I learn in “Parallel Subagent Spawning”?

Multiple Task calls in one response run concurrently. You practise Claude Architect with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.

Do I need any experience to start Claude Architect?

No prior experience is required. Claude Architect on CoddyKit is structured for beginners through advanced learners; this is — lesson 4 of 4, so you can start here or from the beginning and move at your own pace.

How long does the “Parallel Subagent Spawning” lesson take?

Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.

Can I write and run code in this Claude Architect lesson?

Yes. Every Claude Architect lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.

All lessons in this course

  1. Hub-and-Spoke Coordinator Topology
  2. Coordinator Responsibilities
  3. Subagents Don't Inherit History
  4. Parallel Subagent Spawning
← Back to Claude Architect