0Pricing
Claude Architect · 강의

중간 누락 효과

모델은 중간보다 시작과 끝을 더 안정적으로 읽습니다

중간 누락 효과은(는) CoddyKit의 무료 Claude Architect 강의입니다. 이것은 4개 중 3번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 Claude Architect 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. Claude Architect 강의에는 총 4개의 강의가 포함되어 있습니다.

이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.

The Effect, Stated Plainly

When you pack a long prompt into Claude's context window, the model does not attend to every token equally. Information near the start and the end of the context is recalled far more reliably than information buried in the middle.

This is the lost-in-the-middle effect. It is not a bug you can patch — it is a property of how attention behaves over long inputs. As an architect, you design around it.

Why It Matters for Reliability

On the Claude Certified Architect exam, this sits in Domain 5: Context Management & Reliability. The practical risk: a critical instruction, policy rule, or transactional fact that you placed in the middle of a bloated prompt gets silently ignored.

The failure is quiet. The model still produces a fluent answer — it just dropped the constraint you cared about. That makes lost-in-the-middle a reliability problem, not a formatting preference.

Position Your Key Instructions

The simplest mitigation: put your most important instructions where attention is strongest. Lead with them in the system prompt, and restate the critical constraint at the end, right before the model generates.

Remember the API contract: the model keeps no state. You send the FULL messages history every turn. So every turn is a fresh chance — and a fresh risk — for middle content to be under-weighted.

resp = client.messages.create(
    model="claude-sonnet-4-5",
    max_tokens=1024,
    system=(
        "You are a refund agent. CRITICAL RULE: never "
        "process a refund before get_customer returns a "
        "verified ID."
    ),
    # ...long retrieved context goes in the middle...
    messages=messages,
)

Bookend the Critical Constraint

A robust pattern is to bookend: state the rule up front in system, then echo a short reminder of it as the final line of the last user turn. The start anchors it; the end keeps it top-of-mind at generation time.

Keep the echo short. You are not duplicating the whole instruction — just the single decision the model must not forget.

messages = [
    {"role": "user", "content": (
        retrieved_docs
        + "\n\n---\nReminder: cite a source URL for every "
          "claim. If a fact is absent from the docs above, "
          "say so explicitly — do not invent it."
    )},
]

Trim Verbose Tool Output

A major source of middle-bloat is raw tool output. An API or MCP tool can return hundreds of fields when you need three. Dumping all of it pushes your real signal into the low-attention middle.

Trim tool results to the relevant fields before appending them to the message history. Less noise in the middle means the model spends its attention on what matters.

def trim_order(raw: dict) -> dict:
    # keep only what the model needs to reason about
    return {
        "order_id": raw["order_id"],
        "status": raw["status"],
        "total": raw["total"],
        "refundable": raw["refundable"],
    }

# append the trimmed result, not the 200-field blob
messages.append({
    "role": "user",
    "content": [{
        "type": "tool_result",
        "tool_use_id": tu_id,
        "content": json.dumps(trim_order(raw_result)),
    }],
})

Keep Case Facts Verbatim

When context grows, teams reach for progressive summarization. Useful — but it makes numbers, percentages, and dates vague. Combined with lost-in-the-middle, a summarized transactional fact sitting mid-context is doubly fragile.

The fix: pull transactional facts into a separate "case facts" block kept verbatim, outside the summary. Don't let an order total or a refund threshold survive only as a paraphrase in the middle of a summary.

case_facts = (
    "CASE FACTS (verbatim, do not summarize):\n"
    "- order_id: A-4821\n"
    "- order_total: $612.40\n"
    "- refund_policy_cap: $500.00\n"
    "- customer_verified: true"
)

system = case_facts + "\n\n" + rolling_summary

/compact Has the Same Risk

In Claude Code, /compact compresses the conversation to free up context. The benefit is room to keep working; the risk is identical to progressive summarization: numbers and dates become vague.

Before compacting a session that hinges on exact values, capture those values somewhere durable — for example in CLAUDE.md via /memory, which persists across sessions — so they survive compaction rather than dissolving into a fuzzy middle summary.

# In a Claude Code session:
/memory   # write exact build numbers / API contract to CLAUDE.md
/compact  # now safe(r): the hard facts persist outside the summary

Order Retrieved Documents Deliberately

If you inject N retrieved chunks, their order changes what the model uses. The most relevant evidence should sit at the start or the end of the block — not at chunk number 7 of 14.

Two practical moves:

  • Re-rank so the highest-relevance chunks bracket the block.
  • Cut the block down. Fewer, sharper chunks beat a long tail of marginal ones diluting attention.
ranked = rerank(query, chunks)            # best first
top = ranked[:6]                          # cut the long tail
# bracket: strongest at start AND end
ordered = [top[0]] + top[2:] + [top[1]]
context = "\n\n".join(c.text for c in ordered)

Split Work, Don't Dilute Attention

Lost-in-the-middle is one reason single-pass multi-file review dilutes attention. Cramming ten files into one prompt buries the middle files.

Decompose instead: a per-file local pass, then a separate cross-file integration pass. Each pass has a focused, shorter context where nothing important is stranded in the middle.

# Multi-pass review (avoids single-prompt dilution)
for path in changed_files:
    review_local(path)        # focused, short context per file

review_cross_file(changed_files)  # separate integration pass

Subagents Need Explicit Context

Multi-agent systems are hub-and-spoke. Subagents do not inherit the coordinator's conversation history — every fact must be passed explicitly in each subagent prompt.

This intersects with lost-in-the-middle: when the coordinator builds that subagent prompt, the critical instructions still belong at the start and end, and verbose blobs still get trimmed. Context isolation is an opportunity to hand each subagent a clean, well-positioned prompt instead of an inherited mess.

Task(
    description="Summarize Q3 revenue",
    prompt=(
        "TASK: extract Q3 net revenue from the report below.\n"
        + report_text +
        "\n\nOUTPUT: a single number in USD, plus the source "
        "page. If absent, reply 'NOT FOUND'."
    ),
)

A Mental Checklist

Before you ship a long-context prompt, run this checklist:

  • Top: most important instruction in system, stated explicitly.
  • Bottom: short echo of the single critical constraint at the end.
  • Middle: trimmed tool output, re-ranked docs, no raw blobs.
  • Facts: exact numbers/dates kept verbatim in a case-facts block, never only in a summary.
  • Length: if it's huge, split into focused passes.

Position is design. Treat the middle as low-trust real estate.

Quick Check

Apply the lesson to a concrete architecture decision.

Recap

Key takeaways on the lost-in-the-middle effect:

  • Models attend to the start and end more than the middle — treat the middle as low-trust space.
  • Put critical instructions in system and echo the one constraint that matters at the end.
  • Trim verbose tool output and re-rank retrieved docs so signal brackets the context.
  • Keep exact numbers and dates verbatim in a case-facts block; summarization and /compact make them vague.
  • For large work, split into focused passes rather than one diluted prompt.
  • Position mitigates the symptom — but for rules with financial, legal, or safety stakes, enforce with hooks, not placement.

자주 묻는 질문

“중간 누락 효과” 강의는 무료인가요?

네 — “중간 누락 효과” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 Claude Architect 강의 전체를 잠금 해제할 수 있습니다. Claude Architect 강의에는 총 4개의 강의가 포함되어 있습니다.

“중간 누락 효과”에서 뭘 배우나요?

모델은 중간보다 시작과 끝을 더 안정적으로 읽습니다 브라우저에서 직접 실행하는 실습 코드로 Claude Architect을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.

Claude Architect을(를) 시작하는 데 경험이 필요한가요?

사전 경험은 필요하지 않습니다. CoddyKit의 Claude Architect은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 3번째 강의입니다.

“중간 누락 효과” 강의는 얼마나 걸리나요?

대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.

이 Claude Architect 강의에서 코드를 작성하고 실행할 수 있나요?

네. 모든 Claude Architect 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.

이 강의의 모든 강의

  1. 전체 이력이 필요합니다
  2. 점진적 요약의 위험
  3. 중간 누락 효과
  4. 사례 사실 블록 및 출력 잘라내기
← Claude Architect(으)로 돌아가기