Instructor: Pydantic ile Tür Bilgili Çıkarma
Çıkarma başarılı olana kadar yanıtları Pydantic şemanıza göre otomatik olarak yeniden deneyip doğrulaması için OpenAI istemcisini yamalamak üzere instructor kütüphanesini kullanın.
Instructor: Pydantic ile Tür Bilgili Çıkarma, CoddyKit'te ücretsiz bir AI Engineering Academy dersidir. Bu, 4 dersinin 1. dersidir. Aşağıdan dersin tamamını ücretsiz okuyabilir, sonra tarayıcıda yerleşik kod editörü ve 7/24 yapay zeka koçu ile uygulamalı olarak pratik yapabilirsin. Bu, AI Engineering Academy öğrenme yolunun bir parçasıdır ve ilerlemeniz web ve CoddyKit uygulaması arasında senkronize olur. AI Engineering Academy kursu toplamda 4 dersten oluşur.
Bu dersin bazı bölümleri henüz çevrilmemiş olup İngilizce olarak gösterilmektedir.
What Is the Instructor Library?
The instructor library is a thin wrapper around the OpenAI client that makes structured extraction reliable. Instead of hoping the model returns valid JSON, instructor enforces your Pydantic schema and automatically retries if validation fails. It eliminates the need to write custom parsing and retry logic yourself.
Installing Instructor
Install instructor with a single pip command. It requires pydantic v2 and the openai SDK. Once installed, you patch the OpenAI client with instructor.patch() to get the enhanced client that supports the response_model parameter on every call.
pip install instructor openai pydanticPatching the OpenAI Client
Instructor works by patching the standard OpenAI client. Calling instructor.from_openai(client) returns a new client where every chat.completions.create call accepts a response_model keyword argument. The underlying API call is identical — instructor just adds schema enforcement on top.
import instructor
from openai import OpenAI
client = instructor.from_openai(OpenAI())Defining Your Pydantic Schema
Define the data shape you want back from the model as a Pydantic BaseModel. Field names, types, and docstrings are automatically converted into the JSON Schema sent to the model. Use clear, descriptive field names so the model understands what to populate. Add validators for business rules.
from pydantic import BaseModel, Field
from typing import Optional
class PersonExtract(BaseModel):
name: str = Field(description='Full name of the person')
age: Optional[int] = Field(None, description='Age in years if mentioned')
email: Optional[str] = Field(None, description='Email address if present')
company: Optional[str] = Field(None, description='Company or employer')Making an Extraction Call
Pass your Pydantic model class as response_model to the patched client. Instructor constructs a tool call behind the scenes, the model fills in the fields, and instructor deserializes the result into a typed Python object. You get full IDE autocompletion and type safety on the returned data.
result = client.chat.completions.create(
model='gpt-4o-mini',
response_model=PersonExtract,
messages=[
{'role': 'user', 'content': 'Alice Smith, 34, works at Acme Corp. Email: alice@acme.com'}
]
)
print(result.name) # Alice Smith
print(result.email) # alice@acme.comAutomatic Retry on Validation Failure
If the model returns data that fails Pydantic validation, instructor automatically sends the validation error back to the model and asks it to fix its response. You can configure the maximum number of retries with the max_retries parameter. This self-healing loop eliminates most one-off extraction failures without any extra code.
import instructor
from openai import OpenAI
from pydantic import BaseModel, field_validator
client = instructor.from_openai(OpenAI())
class Product(BaseModel):
name: str
price_usd: float
@field_validator('price_usd')
@classmethod
def must_be_positive(cls, v):
if v <= 0:
raise ValueError('Price must be positive')
return v
result = client.chat.completions.create(
model='gpt-4o-mini',
response_model=Product,
max_retries=3,
messages=[{'role': 'user', 'content': 'Widget costs $12.99'}]
)Nested Models for Complex Structures
Instructor handles nested Pydantic models seamlessly. You can define deeply nested schemas with lists, optional sub-objects, and discriminated unions. The model receives the full JSON Schema and must populate all required fields, making it ideal for extracting structured objects like invoices or resumes with multiple sections.
from pydantic import BaseModel
from typing import List
class LineItem(BaseModel):
description: str
quantity: int
unit_price: float
class Invoice(BaseModel):
vendor: str
invoice_number: str
total_amount: float
line_items: List[LineItem]
result = client.chat.completions.create(
model='gpt-4o',
response_model=Invoice,
messages=[{'role': 'user', 'content': invoice_text}]
)Streaming Partial Extractions
For large extraction jobs, instructor supports partial streaming via instructor.Partial[YourModel]. As the model generates tokens, you receive partially populated model instances in real time. This is useful for showing progress in a UI or processing fields as soon as they arrive, rather than waiting for the complete response.
import instructor
from openai import OpenAI
client = instructor.from_openai(OpenAI())
for partial in client.chat.completions.create_partial(
model='gpt-4o-mini',
response_model=PersonExtract,
messages=[{'role': 'user', 'content': long_text}]
):
print(partial.name, partial.email)Extracting Lists of Objects
When you need to extract multiple entities from a single document, wrap your model in List[YourModel]. Instructor handles the JSON array schema and deserializes each element into a typed Python object. This pattern works well for extracting all people mentioned in an article, all transactions in a statement, or all dates in a contract.
from pydantic import BaseModel
from typing import List
class Mention(BaseModel):
entity: str
entity_type: str # PERSON, ORG, DATE, LOCATION
context: str
result = client.chat.completions.create(
model='gpt-4o-mini',
response_model=List[Mention],
messages=[{'role': 'user', 'content': article_text}]
)
for mention in result:
print(f'{mention.entity} ({mention.entity_type})')Choosing the Right Model for Extraction
Not all extractions require GPT-4o. For simple flat schemas with fewer than 10 fields, gpt-4o-mini produces near-identical results at one-tenth the cost. Use GPT-4o for complex nested schemas, long documents, or cases where recall matters. Always benchmark on a sample of your real data before choosing a model for production.
# Cost comparison for 1000 extractions
# GPT-4o-mini: ~$0.002 per call = $2.00 total
# GPT-4o: ~$0.015 per call = $15.00 total
# Test both on 50 samples and compare F1 score
# before committing to the expensive modelLogging and Debugging Extractions
Instructor exposes a hooks system for observability. Register a on_completion callback to log the raw API response, token usage, and number of retries for each extraction. This helps you identify which document types cause the most failures and tune your schemas or prompts accordingly.
import instructor
from openai import OpenAI
client = instructor.from_openai(OpenAI())
@client.on('completion:response')
def log_usage(response):
usage = response.usage
print(f'Tokens: {usage.prompt_tokens}+{usage.completion_tokens}')
result = client.chat.completions.create(
model='gpt-4o-mini',
response_model=PersonExtract,
messages=[{'role': 'user', 'content': text}]
)Quick Check
Test your understanding of the instructor library for typed extraction.
Lesson Recap
In this lesson you learned: instructor patches the OpenAI client to accept a response_model parameter that enforces Pydantic schemas, automatic retry on validation failure makes extraction robust without manual error handling, and nested models and list extraction let you parse complex multi-entity documents into fully typed Python objects. Next up we handle partial and missing data in extracted schemas.
Sıkça Sorulan Sorular
“Instructor: Pydantic ile Tür Bilgili Çıkarma” dersi ücretsiz mi?
Evet — “Instructor: Pydantic ile Tür Bilgili Çıkarma” dersin tüm metni burada web'de ücretsiz olarak okunabilir. Etkileşimli olarak pratik yapmak (yerleşik kod editörü ve 7/24 yapay zeka koçu) ve AI Engineering Academy kursunun geri kalanını açmak için CoddyKit PRO'ya yükselt. AI Engineering Academy kursu toplamda 4 dersten oluşur.
“Instructor: Pydantic ile Tür Bilgili Çıkarma” dersinde ne öğreneceğim?
Çıkarma başarılı olana kadar yanıtları Pydantic şemanıza göre otomatik olarak yeniden deneyip doğrulaması için OpenAI istemcisini yamalamak üzere instructor kütüphanesini kullanın. AI Engineering Academy ile uygulamalı kodu tarayıcıda doğrudan çalıştırarak pratik yaparsın ve 7/24 yapay zeka koçu dersi çalışırken sorularını yanıtlar.
AI Engineering Academy öğrenmeye başlamak için deneyim gerekli mi?
Önceden deneyim gerekmez. CoddyKit'te AI Engineering Academy, başlangıçtan ileri seviyeye kadar yapılandırıldığı için buradan başlayabilir veya başından başlayıp kendi hızında ilerleme yapabilirsin. Bu, 4 dersinin 1. dersidir.
“Instructor: Pydantic ile Tür Bilgili Çıkarma” dersi ne kadar sürer?
Çoğu CoddyKit dersi yaklaşık 5–10 dakika sürer. Her biri kısa ve etkileşimli olduğu için sabit ilerleme yaparsın ve web ile uygulama arasında tam olarak bıraktığın yerden devam edebilirsin.
Bu AI Engineering Academy dersinde kod yazıp çalıştırabilir miyim?
Evet. Her AI Engineering Academy dersi yerleşik bir kod editörü içerir, bu sayede tarayıcıda gerçek kod yazıp çalıştırabilir ve anlık yapay zeka geri bildirimi alırsın — yerel kurulum gerekli değildir.
Bu kursun tüm dersleri
- Instructor: Pydantic ile Tür Bilgili Çıkarma
- Kısmi ve Eksik Verileri İşleme
- Eşzamansız İşleme ve Kuyruklarla Toplu İşleme
- Şema Gelişimi ve Geriye Dönük Uyumluluk