Instructor: Pydantic을 활용한 타입 지정 추출
instructor 라이브러리를 사용해 OpenAI 클라이언트를 수정하고, 추출이 성공할 때까지 Pydantic 스키마에 따라 응답을 자동으로 재시도하고 검증하게 합니다.
Instructor: Pydantic을 활용한 타입 지정 추출은(는) CoddyKit의 무료 AI Engineering Academy 강의입니다. 이것은 4개 중 1번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 AI Engineering Academy 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. AI Engineering Academy 강의에는 총 4개의 강의가 포함되어 있습니다.
이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.
What Is the Instructor Library?
The instructor library is a thin wrapper around the OpenAI client that makes structured extraction reliable. Instead of hoping the model returns valid JSON, instructor enforces your Pydantic schema and automatically retries if validation fails. It eliminates the need to write custom parsing and retry logic yourself.
Installing Instructor
Install instructor with a single pip command. It requires pydantic v2 and the openai SDK. Once installed, you patch the OpenAI client with instructor.patch() to get the enhanced client that supports the response_model parameter on every call.
pip install instructor openai pydanticPatching the OpenAI Client
Instructor works by patching the standard OpenAI client. Calling instructor.from_openai(client) returns a new client where every chat.completions.create call accepts a response_model keyword argument. The underlying API call is identical — instructor just adds schema enforcement on top.
import instructor
from openai import OpenAI
client = instructor.from_openai(OpenAI())Defining Your Pydantic Schema
Define the data shape you want back from the model as a Pydantic BaseModel. Field names, types, and docstrings are automatically converted into the JSON Schema sent to the model. Use clear, descriptive field names so the model understands what to populate. Add validators for business rules.
from pydantic import BaseModel, Field
from typing import Optional
class PersonExtract(BaseModel):
name: str = Field(description='Full name of the person')
age: Optional[int] = Field(None, description='Age in years if mentioned')
email: Optional[str] = Field(None, description='Email address if present')
company: Optional[str] = Field(None, description='Company or employer')Making an Extraction Call
Pass your Pydantic model class as response_model to the patched client. Instructor constructs a tool call behind the scenes, the model fills in the fields, and instructor deserializes the result into a typed Python object. You get full IDE autocompletion and type safety on the returned data.
result = client.chat.completions.create(
model='gpt-4o-mini',
response_model=PersonExtract,
messages=[
{'role': 'user', 'content': 'Alice Smith, 34, works at Acme Corp. Email: alice@acme.com'}
]
)
print(result.name) # Alice Smith
print(result.email) # alice@acme.comAutomatic Retry on Validation Failure
If the model returns data that fails Pydantic validation, instructor automatically sends the validation error back to the model and asks it to fix its response. You can configure the maximum number of retries with the max_retries parameter. This self-healing loop eliminates most one-off extraction failures without any extra code.
import instructor
from openai import OpenAI
from pydantic import BaseModel, field_validator
client = instructor.from_openai(OpenAI())
class Product(BaseModel):
name: str
price_usd: float
@field_validator('price_usd')
@classmethod
def must_be_positive(cls, v):
if v <= 0:
raise ValueError('Price must be positive')
return v
result = client.chat.completions.create(
model='gpt-4o-mini',
response_model=Product,
max_retries=3,
messages=[{'role': 'user', 'content': 'Widget costs $12.99'}]
)Nested Models for Complex Structures
Instructor handles nested Pydantic models seamlessly. You can define deeply nested schemas with lists, optional sub-objects, and discriminated unions. The model receives the full JSON Schema and must populate all required fields, making it ideal for extracting structured objects like invoices or resumes with multiple sections.
from pydantic import BaseModel
from typing import List
class LineItem(BaseModel):
description: str
quantity: int
unit_price: float
class Invoice(BaseModel):
vendor: str
invoice_number: str
total_amount: float
line_items: List[LineItem]
result = client.chat.completions.create(
model='gpt-4o',
response_model=Invoice,
messages=[{'role': 'user', 'content': invoice_text}]
)Streaming Partial Extractions
For large extraction jobs, instructor supports partial streaming via instructor.Partial[YourModel]. As the model generates tokens, you receive partially populated model instances in real time. This is useful for showing progress in a UI or processing fields as soon as they arrive, rather than waiting for the complete response.
import instructor
from openai import OpenAI
client = instructor.from_openai(OpenAI())
for partial in client.chat.completions.create_partial(
model='gpt-4o-mini',
response_model=PersonExtract,
messages=[{'role': 'user', 'content': long_text}]
):
print(partial.name, partial.email)Extracting Lists of Objects
When you need to extract multiple entities from a single document, wrap your model in List[YourModel]. Instructor handles the JSON array schema and deserializes each element into a typed Python object. This pattern works well for extracting all people mentioned in an article, all transactions in a statement, or all dates in a contract.
from pydantic import BaseModel
from typing import List
class Mention(BaseModel):
entity: str
entity_type: str # PERSON, ORG, DATE, LOCATION
context: str
result = client.chat.completions.create(
model='gpt-4o-mini',
response_model=List[Mention],
messages=[{'role': 'user', 'content': article_text}]
)
for mention in result:
print(f'{mention.entity} ({mention.entity_type})')Choosing the Right Model for Extraction
Not all extractions require GPT-4o. For simple flat schemas with fewer than 10 fields, gpt-4o-mini produces near-identical results at one-tenth the cost. Use GPT-4o for complex nested schemas, long documents, or cases where recall matters. Always benchmark on a sample of your real data before choosing a model for production.
# Cost comparison for 1000 extractions
# GPT-4o-mini: ~$0.002 per call = $2.00 total
# GPT-4o: ~$0.015 per call = $15.00 total
# Test both on 50 samples and compare F1 score
# before committing to the expensive modelLogging and Debugging Extractions
Instructor exposes a hooks system for observability. Register a on_completion callback to log the raw API response, token usage, and number of retries for each extraction. This helps you identify which document types cause the most failures and tune your schemas or prompts accordingly.
import instructor
from openai import OpenAI
client = instructor.from_openai(OpenAI())
@client.on('completion:response')
def log_usage(response):
usage = response.usage
print(f'Tokens: {usage.prompt_tokens}+{usage.completion_tokens}')
result = client.chat.completions.create(
model='gpt-4o-mini',
response_model=PersonExtract,
messages=[{'role': 'user', 'content': text}]
)Quick Check
Test your understanding of the instructor library for typed extraction.
Lesson Recap
In this lesson you learned: instructor patches the OpenAI client to accept a response_model parameter that enforces Pydantic schemas, automatic retry on validation failure makes extraction robust without manual error handling, and nested models and list extraction let you parse complex multi-entity documents into fully typed Python objects. Next up we handle partial and missing data in extracted schemas.
자주 묻는 질문
“Instructor: Pydantic을 활용한 타입 지정 추출” 강의는 무료인가요?
네 — “Instructor: Pydantic을 활용한 타입 지정 추출” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 AI Engineering Academy 강의 전체를 잠금 해제할 수 있습니다. AI Engineering Academy 강의에는 총 4개의 강의가 포함되어 있습니다.
“Instructor: Pydantic을 활용한 타입 지정 추출”에서 뭘 배우나요?
instructor 라이브러리를 사용해 OpenAI 클라이언트를 수정하고, 추출이 성공할 때까지 Pydantic 스키마에 따라 응답을 자동으로 재시도하고 검증하게 합니다. 브라우저에서 직접 실행하는 실습 코드로 AI Engineering Academy을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.
AI Engineering Academy을(를) 시작하는 데 경험이 필요한가요?
사전 경험은 필요하지 않습니다. CoddyKit의 AI Engineering Academy은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 1번째 강의입니다.
“Instructor: Pydantic을 활용한 타입 지정 추출” 강의는 얼마나 걸리나요?
대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.
이 AI Engineering Academy 강의에서 코드를 작성하고 실행할 수 있나요?
네. 모든 AI Engineering Academy 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.
이 강의의 모든 강의
- Instructor: Pydantic을 활용한 타입 지정 추출
- 부분 데이터와 누락 데이터 처리
- 비동기 처리와 대기열을 활용한 일괄 처리
- 스키마 진화와 이전 버전 호환성