CI/CD 및 구조화된 추출
헤드리스 모드, JSON 출력, 스키마 및 검증 루프를 다룹니다
CI/CD 및 구조화된 추출은(는) CoddyKit의 무료 Claude Architect 강의입니다. 이것은 4개 중 3번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 Claude Architect 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. Claude Architect 강의에는 총 4개의 강의가 포함되어 있습니다.
이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.
Two Scenarios, One Theme
This lesson fuses two exam scenarios that share a single backbone: determinism under automation. Scenario 5 (Claude Code for CI/CD) and Scenario 6 (Structured Data Extraction) both ask the same question — how do you get machine-parseable, trustworthy output with no human in the loop?
The answer is the same in both worlds:
- Headless / non-interactive execution so a pipeline can drive Claude.
- JSON output with a schema so downstream code can parse results.
- Validation loops that catch and repair structural errors before they propagate.
Master this and you cover a meaningful slice of D3 (Config & Workflows) and D4 (Structured Output) — together 40% of the exam.
Headless Mode in CI/CD
A pipeline has no TTY and no human to answer prompts. Claude Code must run non-interactively. The flag for this is -p (also written --print): it runs a single prompt, prints the result, and exits.
Pair it with --output-format json so the pipeline gets a structured envelope instead of free prose. Without these two flags, the command hangs waiting for interactive input and the build stalls.
-p/--print= non-interactive, required in pipelines.--output-format json= parseable result (optionally with a schema).
# Headless review step in a CI job
claude -p "Review the staged diff for correctness bugs only. \
Flag a comment ONLY when it contradicts the code." \
--output-format json \
> review.json
# Pipeline now parses review.json deterministically
jq '.result' review.jsonThe Isolated Review Session
A subtle but heavily-tested rule: in CI, review code in an ISOLATED session, separate from the session that generated it.
Why? The generating session retains its own reasoning and is biased toward defending its work — same-session self-review is a top anti-pattern. A fresh, independent instance has no attachment to the output and challenges it honestly.
This mirrors the extraction-world rule that independent/fresh-instance review beats same-session self-review. The author won't fight its own conclusions; a clean reviewer will.
Minimizing False Positives
A CI reviewer that cries wolf gets ignored. The goal in Scenario 5 is to minimize false positives so developers trust the gate.
The lever is explicit criteria, not vague pleas. "Be more precise" changes nothing. "Flag a comment ONLY when it contradicts the code" gives the model a sharp, testable boundary.
On re-runs, don't re-litigate everything: include the prior results and report only new or still-unfixed issues. This keeps the signal clean across iterations.
claude -p "You are reviewing a re-run. Here are the prior findings:
$(cat prev_findings.json)
Report ONLY issues that are new or remain unfixed.
Flag a defect ONLY when the code's behavior contradicts its stated intent.
Ignore style and subjective preferences." \
--output-format json > findings.jsonBlocking Checks vs the Batch API
A classic distractor: "use the Batch API to cut CI costs by 50%". Wrong for a pre-merge gate.
Message Batches are 50% cheaper with up to a 24h window, but they have NO latency SLA and do NOT support multi-turn tool calling. A blocking, time-sensitive check (like a PR gate) cannot wait an unbounded amount of time.
- Use Batch for non-blocking work: overnight reports, large audits, nightly extraction runs.
- Do NOT use Batch for pre-merge / blocking / interactive checks.
For batch jobs, custom_id correlates each request to its result, and you re-submit only the failures.
Structured Output: Tool Use + JSON Schema
Now Scenario 6. To extract data reliably, do not ask for JSON in prose and hope. Use tool_use with a JSON Schema: this eliminates syntax errors and enforces required fields.
Set tool_choice to "any" to guarantee the model calls a tool (i.e. produces structured output) rather than free text. Use {"type":"tool","name":"X"} to force one specific extraction tool.
extract_tool = {
"name": "extract_invoice",
"description": "Extract structured fields from an invoice document.",
"input_schema": {
"type": "object",
"properties": {
"invoice_number": {"type": "string"},
"line_items": {"type": "array", "items": {"type": "object"}},
"stated_total": {"type": "number"},
},
"required": ["invoice_number"],
},
}
resp = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=1024,
tools=[extract_tool],
tool_choice={"type": "any"}, # MUST call a tool -> structured output
messages=[{"role": "user", "content": document_text}],
)Required Fields: The Fabrication Trap
The single most-tested schema rule: mark a field required ONLY if it is always present in the source.
If you require a field that may be absent, the model has no legal way to satisfy the schema except to fabricate a value. You traded a missing field for a hallucinated one — far worse for a data pipeline.
For optional or open-ended values, prefer an enum with an "other" value plus a free-text detail field. This keeps the output structured while staying extensible for cases you didn't enumerate.
"document_type": {
"type": "string",
"enum": ["invoice", "receipt", "purchase_order", "other"]
},
"document_type_detail": {
"type": "string",
"description": "Free text. Required only when document_type is 'other'."
}
# 'required' lists ONLY fields guaranteed to appear -> no fabrication pressureThe Validation / Retry Loop
A schema constrains shape, not correctness. Wrap extraction in a validation loop (Pydantic-style validation is the canonical pattern).
When validation fails, use retry-with-feedback: send the model three things together — the original document, the wrong output it produced, and the exact validation error. This is precise enough to fix format, structural, and arithmetic mistakes.
Critical boundary: retry fixes FORMAT errors. It does NOT help when the information is simply ABSENT from the source — re-asking a document for data it never contained just invites a fabricated answer.
from pydantic import BaseModel, ValidationError
def extract_with_retry(doc, max_attempts=2):
history = [{"role": "user", "content": doc}]
for _ in range(max_attempts): # cap = safety net, not primary stop
out = call_extractor(history)
try:
return InvoiceModel.model_validate(out)
except ValidationError as e:
history.append({"role": "assistant", "content": str(out)})
history.append({"role": "user", "content":
f"Validation failed: {e}. Re-extract from the SAME document. "
f"If a field is not present in the source, omit it; do not invent it."})
raise ExtractionError("unresolved after retries")Self-Correction: Detecting Discrepancies
How do you catch a wrong arithmetic value the schema happily accepted? Build the check into the extraction itself.
The pattern: extract both the calculated_total (sum the line items) and the stated_total (the figure printed on the document). Then compare them in code.
If they diverge, you've detected a discrepancy deterministically — either a document error or an extraction error — and can flag, retry, or escalate. One number can lie silently; two numbers expose the lie.
# Schema asks for BOTH so code can self-correct
"calculated_total": {"type": "number",
"description": "Sum of all line_items, computed by you."},
"stated_total": {"type": "number",
"description": "The total figure printed on the document."}
# Downstream deterministic check
if abs(result.calculated_total - result.stated_total) > 0.01:
flag_for_review(result, reason="total mismatch")Provenance: Claim to Source
Extraction without provenance is unauditable. For every extracted claim, keep a claim-to-source mapping: the source document name or URL, the supporting quote, and the publication date.
When two sources give conflicting figures, do not arbitrarily pick one — annotate the conflict. Often the dates resolve the apparent contradiction (one figure is simply newer).
And render by content type: tables for financials, prose for narrative, lists for technical findings. Provenance plus correct rendering is what makes the output trustworthy enough to automate on.
"fields": [{
"name": "annual_revenue",
"value": "4.2B USD",
"source_doc": "FY24-10K.pdf",
"quote": "Total revenue was $4.2 billion in fiscal 2024.",
"published": "2024-03-01"
}]
# Conflicting value from an older filing? Annotate, don't overwrite.Metrics: Aggregate Accuracy Lies
Before you let an extraction pipeline run unattended, validate it honestly. A headline like "97% accurate" can hide poor performance on one specific document type or one field.
The disciplined approach:
- Stratified random sampling across document types — not a convenience sample.
- Field-level confidence, calibrated on a labeled validation set, before automating.
Aggregate-only metrics are a known anti-pattern. A 97% average with 40% accuracy on tax forms is a production incident waiting to happen.
Exam Scenario: The Pre-Merge Gate
Apply the full picture to a realistic exam decision.
Recap: Deterministic Output Under Automation
Key takeaways for CI/CD and structured extraction:
- Headless:
-p/--print+--output-format jsonfor non-interactive, parseable pipeline runs. - Isolated review beats same-session self-review; explicit criteria minimize false positives; on re-runs report only new/unfixed issues.
- Batch API = 50% cheaper, no latency SLA, no multi-turn tools — for overnight jobs, NEVER blocking checks.
- Structured output: tool_use + JSON Schema;
tool_choice:"any"guarantees a structured call. - Required only if always present — requiring an absent field forces fabrication; use enum + 'other' + detail for extensibility.
- Retry-with-feedback (original doc + wrong output + exact error) fixes FORMAT errors, not ABSENT data.
- Self-correct with calculated vs stated totals; keep provenance (source, quote, date) and annotate conflicts.
- Validate with stratified sampling + field-level confidence — aggregate accuracy hides weak spots.
자주 묻는 질문
“CI/CD 및 구조화된 추출” 강의는 무료인가요?
네 — “CI/CD 및 구조화된 추출” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 Claude Architect 강의 전체를 잠금 해제할 수 있습니다. Claude Architect 강의에는 총 4개의 강의가 포함되어 있습니다.
“CI/CD 및 구조화된 추출”에서 뭘 배우나요?
헤드리스 모드, JSON 출력, 스키마 및 검증 루프를 다룹니다 브라우저에서 직접 실행하는 실습 코드로 Claude Architect을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.
Claude Architect을(를) 시작하는 데 경험이 필요한가요?
사전 경험은 필요하지 않습니다. CoddyKit의 Claude Architect은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 3번째 강의입니다.
“CI/CD 및 구조화된 추출” 강의는 얼마나 걸리나요?
대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.
이 Claude Architect 강의에서 코드를 작성하고 실행할 수 있나요?
네. 모든 Claude Architect 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.
이 강의의 모든 강의
- 지원 에이전트 및 다중 에이전트 조사
- 코드 생성 및 개발자 생산성
- CI/CD 및 구조화된 추출
- 대화형 패턴 및 에이전트형 도구