Параллельный запуск подагентов
Несколько вызовов Task в одном ответе выполняются одновременно
«Параллельный запуск подагентов» — бесплатный урок Claude Architect на CoddyKit. Это урок 4 из 4. Ты можешь прочитать весь урок бесплатно ниже — а потом практиковать его прямо в браузере с встроенным редактором кода и ИИ-репетитором 24/7. Это часть пути обучения Claude Architect, и твой прогресс синхронизируется между веб-версией и приложением CoddyKit. Курс Claude Architect содержит 4 уроков всего.
Части этого урока еще не переведены и отображаются на английском.
Why Parallelism Matters
In a hub-and-spoke multi-agent system, a coordinator decomposes a task, delegates pieces to subagents, then aggregates the results. The key performance lever: multiple Task calls emitted in one response run concurrently.
If a research job needs three independent lookups, you don't have to do them one after another. Spawn all three in a single turn and they execute in parallel, collapsing three round-trips into one wave.
The Parallel Rule
The rule is precise: every Task call the model emits in the same assistant response is dispatched at the same time. There is no special API flag — concurrency is a property of how many Task calls share one response.
- Three Task calls in one response → three subagents run in parallel.
- One Task call, wait for the result, then another Task call next turn → sequential.
So parallelism is a decomposition decision, not a configuration toggle.
Coordinator Must Allow Task
A coordinator can only spawn subagents if its allowedTools includes "Task". Without it, the model has no mechanism to delegate and will try to do everything itself.
This is part of least-privilege design: only the coordinator gets Task; leaf subagents usually do not, so they can't recursively spawn more agents.
from claude_agent_sdk import AgentDefinition
coordinator = AgentDefinition(
name="research_coordinator",
description="Decomposes a research question and delegates to subagents in parallel.",
system_prompt="Break the question into independent sub-questions and spawn one subagent per topic IN A SINGLE RESPONSE.",
allowed_tools=["Task"], # required to spawn subagents
)Subagents Don't Inherit History
This is the single most-tested fact about subagents: a subagent does NOT inherit the coordinator's conversation history. It starts with a clean context.
Therefore all context a subagent needs must be passed explicitly in its prompt — the question, relevant constraints, prior findings, output format. If you assume the subagent "already knows" what the coordinator discussed, it will hallucinate or under-perform.
Self-Contained Subagent Prompts
Because context isn't inherited, each parallel Task prompt should be fully self-contained. Notice how each spawned subagent below carries its own topic and the same explicit instructions — nothing is left implicit.
# Coordinator emits these THREE Task calls in ONE response -> they run in parallel
tasks = [
Task(agent="researcher",
prompt="Topic: 2024 EV battery cost trends.\n"
"Return: 3 findings, each with source URL + publication date."),
Task(agent="researcher",
prompt="Topic: solid-state battery timelines.\n"
"Return: 3 findings, each with source URL + publication date."),
Task(agent="researcher",
prompt="Topic: lithium supply constraints 2025.\n"
"Return: 3 findings, each with source URL + publication date."),
]When Parallel Is Correct
Parallel spawning is the right move when sub-tasks are independent — none needs the output of another.
- Multi-topic research: each topic is its own lane.
- Per-file local code review: review each file independently in parallel, then a separate cross-file integration pass.
- Fan-out lookups: several catalogs or APIs queried at once.
If task B needs A's result, that's a sequential dependency — use a fixed pipeline / prompt chaining instead, not parallel spawning.
Parallel vs. Sequential Decomposition
Match the decomposition pattern to the work:
- Parallel Task fan-out → independent subtasks, latency-sensitive, results merged at the end.
- Fixed pipeline / prompt chaining → known sequential steps where each feeds the next.
- Adaptive decomposition → open-ended investigation where the next step depends on what you just found.
Forcing inherently sequential steps into one parallel wave produces agents working with missing inputs — a correctness bug, not a speedup.
# Sequential: step 2 NEEDS step 1's output -> chain, do NOT parallelize
# 1) fetch the schema, THEN 2) extract rows against that schema
schema = Task(agent="schema_reader", prompt="Return the orders table schema.")
# ...wait for result, then next turn:
extract = Task(agent="extractor",
prompt=f"Using this schema:\n{schema}\nExtract Q1 orders.")Aggregating Parallel Results
After a parallel wave returns, the coordinator's job is to aggregate: merge the findings, deduplicate, and resolve conflicts. Don't arbitrarily pick one number when two subagents disagree — annotate the conflict (dates often resolve apparent contradictions).
The coordinator also routes follow-ups: if one lane came back thin, it can spawn a targeted second wave.
# After the parallel wave, the coordinator receives all three result blocks
# and synthesizes. Keep claim -> source mappings for provenance.
synthesis_prompt = (
"You have results from 3 parallel researchers below.\n"
"Merge into one report. For each claim keep its source URL + date.\n"
"If two sources conflict, annotate both rather than picking one."
)Partial Failure in a Parallel Wave
When you fan out, one lane may fail while others succeed. Do not abort the whole workflow on a single failure, and never silently suppress it.
- Recover transient faults locally inside the subagent (retry).
- Escalate non-recoverable failures with partial results and structured context: failure type, attempted query, what did succeed.
The coordinator should annotate coverage gaps ("topic 2 unavailable") rather than presenting incomplete output as complete.
# A subagent returns STRUCTURED context on failure, not a silent drop
{
"isError": True,
"errorCategory": "transient", # transient|validation|business|permission
"isRetryable": True,
"message": "Search API timed out",
"attempted_query": "solid-state battery timelines",
"partial_results": ["1 finding retrieved before timeout"],
}Least-Privilege Subagents
Each subagent's allowed_tools should be scoped to its role — least privilege. A researcher needs read/search tools; it does not need Task (no recursive spawning) and it does not need write or refund tools.
Remember the tool-count guideline: 4-5 tools per agent is optimal; 18+ degrades selection reliability. Narrow, role-scoped subagents pick the right tool far more reliably than one overloaded mega-agent.
researcher = AgentDefinition(
name="researcher",
description="Investigates ONE topic and returns findings with citations.",
system_prompt="Research the given topic. Return findings with source URL + date.",
allowed_tools=["WebSearch", "WebFetch", "Read"], # no Task: cannot re-spawn
)The Agentic Loop Still Governs
Parallelism doesn't change the control flow. The coordinator still runs the standard agentic loop: send request → inspect stop_reason → if tool_use, run the tools (the parallel subagents) and append their results to history → repeat until end_turn.
Terminate on stop_reason, never by parsing text for words like "done". Iteration caps are a safety net, not the primary stop mechanism. Whether one Task or three ran, the loop logic is identical.
Quick Check
A coordinator must research three independent topics as fast as possible. Which approach correctly spawns the subagents in parallel?
Recap
Key takeaways for parallel subagent spawning:
- Multiple Task calls in one response run in parallel — there's no flag, it's a decomposition choice.
- The coordinator's
allowedToolsmust include"Task". - Subagents do not inherit history — pass all context explicitly in each self-contained prompt.
- Use parallel for independent subtasks; use a pipeline/chaining for sequential dependencies.
- On partial failure: recover transient faults locally, escalate non-recoverable ones with partial results — never abort the whole workflow or suppress silently.
- Scope subagents with least privilege (4-5 tools), and keep terminating on
stop_reason, never on parsed text.
Часто задаваемые вопросы
Урок «Параллельный запуск подагентов» бесплатный?
Да — полный текст урока «Параллельный запуск подагентов» бесплатно доступен здесь в веб-версии. Чтобы практиковать его интерактивно (встроенный редактор кода и ИИ-репетитор 24/7) и разблокировать остальной курс Claude Architect, подпишись на CoddyKit PRO. Курс Claude Architect содержит 4 уроков всего.
Чему я научусь в уроке «Параллельный запуск подагентов»?
Несколько вызовов Task в одном ответе выполняются одновременно Ты практикуешь Claude Architect с помощью реального кода, который запускаешь прямо в браузере, и ИИ-репетитор 24/7 отвечает на твои вопросы во время урока.
Нужен ли мне опыт, чтобы начать Claude Architect?
Предыдущий опыт не требуется. Claude Architect на CoddyKit структурирован для всех уровней — от новичков до продвинутых, поэтому ты можешь начать отсюда или с самого начала и учиться в своем темпе. Это урок 4 из 4.
Сколько времени занимает урок «Параллельный запуск подагентов»?
Большинство уроков CoddyKit занимают около 5–10 минут. Каждый из них компактный и интерактивный, поэтому ты постоянно делаешь прогресс и продолжаешь с того же места в веб-версии и приложении.
Можно ли писать и запускать код в этом уроке Claude Architect?
Да. Каждый урок Claude Architect включает встроенный редактор кода, поэтому ты пишешь и запускаешь реальный код прямо в браузере и получаешь моментальную обратную связь от AI — локальная установка не требуется.
Все уроки этого курса
- Топология «центр и лучи» для координатора
- Обязанности координатора
- Подагенты не наследуют историю
- Параллельный запуск подагентов