JSON 스키마 설계
필요한 형태와 정확히 일치하도록 출력을 구성합니다
JSON 스키마 설계은(는) CoddyKit의 무료 Claude Architect 강의입니다. 이것은 4개 중 2번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 Claude Architect 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. Claude Architect 강의에는 총 4개의 강의가 포함되어 있습니다.
이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.
Why Schema-Shaped Output
When you need Claude's answer in a precise structure, don't parse free text and hope. Pair tool_use with a JSON Schema: Claude fills a tool's input_schema and the API guarantees valid JSON with your required fields present.
This eliminates two whole classes of failure: syntax errors (missing commas, unescaped quotes) and missing fields. The schema IS the contract — design it well and downstream code never has to defend against malformed shapes.
The Tool Is the Schema
A structured-output "tool" doesn't have to call anything. It's just a named container whose input_schema describes the shape you want back. You define it, then read what Claude put in the tool call.
Give the tool a clear name and description — these still drive selection — but the real work is in the schema's properties and required list.
extract_invoice = {
"name": "extract_invoice",
"description": "Record the structured fields parsed from an invoice document.",
"input_schema": {
"type": "object",
"properties": {
"invoice_number": {"type": "string"},
"total": {"type": "number"},
},
"required": ["invoice_number", "total"],
},
}Force the Structure with tool_choice
If you want guaranteed structured output, don't leave it to chance. Set tool_choice to force a tool call:
"auto"— model picks text or a tool"any"— model MUST call some tool (guarantees structured output){"type":"tool","name":"X"}— force one specific tool
For single-schema extraction, forcing the exact tool by name is the cleanest path to a deterministic shape.
resp = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=1024,
tools=[extract_invoice],
tool_choice={"type": "tool", "name": "extract_invoice"},
messages=[{"role": "user", "content": invoice_text}],
)Required Means Always Present
The single most important schema rule: mark a field required ONLY if it is always present in the source. Never require a field that may be absent.
Why? A required field forces the model to emit a value. If the data isn't there, Claude will fabricate one to satisfy the contract. Optional-but-absent is honest; required-but-missing breeds hallucination.
Optional Fields, Done Right
For fields that may or may not appear — a discount line, a secondary contact, a due date — leave them OUT of required. Describe them clearly so Claude only fills them when the data genuinely exists.
A good description tells the model the input format and the absence rule, so it omits rather than invents.
"properties": {
"invoice_number": {"type": "string"},
"total": {"type": "number"},
"due_date": {
"type": "string",
"description": "ISO 8601 date (YYYY-MM-DD). Omit entirely if no due date is stated."
},
},
"required": ["invoice_number", "total"]Enums Constrain the Output
When a field has a fixed vocabulary — status, category, priority — use an enum. This collapses messy free text ("paid", "PAID", "settled") into one canonical value your code can switch on.
Enums also reduce hallucination: the model must choose from the listed set instead of inventing a label.
"status": {
"type": "string",
"enum": ["draft", "sent", "paid", "overdue", "void"],
"description": "Current invoice status."
}Designing Enums for Extensibility
Rigid enums break when reality grows a new case. The architect's pattern: add an "other" value to the enum AND a free-text detail field to capture what "other" actually was.
Now your schema stays valid for unforeseen inputs, you don't lose information, and you can mine the detail field to decide if a new enum value is warranted.
"category": {
"type": "string",
"enum": ["hardware", "software", "services", "other"]
},
"category_detail": {
"type": "string",
"description": "If category is 'other', describe it here. Omit otherwise."
}Descriptions Do the Teaching
Field descriptions are mini-prompts. Vague keys produce vague output. Spell out the format, give an example, and state edge-case handling right in the schema.
This is the structured-output equivalent of explicit criteria beating vague instructions: "ISO 8601 date, omit if absent" beats a bare due_date: string every time.
"line_items": {
"type": "array",
"description": "One object per billed line. Empty array if none.",
"items": {
"type": "object",
"properties": {
"sku": {"type": "string", "description": "e.g. 'ABC-1024'"},
"qty": {"type": "integer"},
"unit_price": {"type": "number"}
},
"required": ["qty", "unit_price"]
}
}Build in Self-Verification
Great schemas help you catch errors. To verify arithmetic, extract BOTH a calculated and a stated value, then compare them in code.
For example, capture stated_total (printed on the doc) alongside the line items you can sum yourself. A mismatch flags an extraction or document error before it propagates downstream.
"stated_total": {
"type": "number",
"description": "The grand total exactly as printed on the invoice."
}
# In code:
# calc = sum(li['qty'] * li['unit_price'] for li in items)
# if abs(calc - data['stated_total']) > 0.01: flag_discrepancy()Validate, Then Retry with Feedback
The schema guarantees JSON shape, not business correctness. Layer Pydantic-style validation on top. When it fails on a format / structural / arithmetic error, retry — but feed the model what went wrong.
Send the original document, the wrong output, and the exact validation error. That's retry-with-feedback. Note: retrying does NOT help when information is simply absent from the source — no amount of re-asking conjures missing data.
messages = [
{"role": "user", "content": original_doc},
{"role": "assistant", "content": wrong_output},
{"role": "user", "content":
f"Validation failed: {error}. Re-emit corrected JSON."},
]
# Retry fixes format/arithmetic bugs, not missing facts.Capture Provenance in the Schema
For extraction you'll defend later, design fields that preserve provenance: where each claim came from. Keep claim-to-source mappings — source name, quote, page or date.
This turns a black-box extraction into an auditable one, and lets you annotate conflicting values (often a date difference) instead of silently picking one.
"items": {
"type": "array",
"items": {
"type": "object",
"properties": {
"value": {"type": "string"},
"source_quote": {"type": "string",
"description": "Verbatim text supporting this value."},
"source_page": {"type": "integer"}
},
"required": ["value", "source_quote"]
}
}Quick Check: Optional Field
You're extracting purchase orders. Some POs list a discount_code, but most do not. How should the schema treat discount_code?
Recap: Shaping the Output
Key takeaways for designing a JSON Schema:
- tool_use + JSON Schema kills syntax errors and enforces required fields.
- Force shape with
tool_choice:"any"for some tool,{"type":"tool","name":"X"}for a specific one. - Mark
requiredONLY for always-present fields — requiring an absent field causes fabrication. - Use enums for fixed vocabularies; add
"other"+ a detail field for extensibility. - Descriptions teach format, examples, and edge cases.
- Extract calculated AND stated values to self-verify; validate, then retry-with-feedback for format errors (not for absent info).
- Capture provenance for auditable, defensible extraction.
자주 묻는 질문
“JSON 스키마 설계” 강의는 무료인가요?
네 — “JSON 스키마 설계” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 Claude Architect 강의 전체를 잠금 해제할 수 있습니다. Claude Architect 강의에는 총 4개의 강의가 포함되어 있습니다.
“JSON 스키마 설계”에서 뭘 배우나요?
필요한 형태와 정확히 일치하도록 출력을 구성합니다 브라우저에서 직접 실행하는 실습 코드로 Claude Architect을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.
Claude Architect을(를) 시작하는 데 경험이 필요한가요?
사전 경험은 필요하지 않습니다. CoddyKit의 Claude Architect은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 2번째 강의입니다.
“JSON 스키마 설계” 강의는 얼마나 걸리나요?
대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.
이 Claude Architect 강의에서 코드를 작성하고 실행할 수 있나요?
네. 모든 Claude Architect 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.