Modalità JSON e response_format
Abiliterà la modalità JSON nell'API OpenAI, creerà prompt che producano costantemente JSON valido e gestirà i casi in cui il modello riesca comunque a rompere il formato.
Modalità JSON e response_format è una lezione AI Engineering Academy gratuita su CoddyKit. Questa è la lezione 1 di 4. Puoi leggere la lezione completa qui gratuitamente — poi esercitati direttamente nel browser con un editor di codice integrato e un tutor IA disponibile 24/7. Fa parte del percorso di apprendimento AI Engineering Academy, e i tuoi progressi si sincronizzano tra il web e l'app CoddyKit. Il corso AI Engineering Academy include 4 lezioni in totale.
Parti di questa lezione non sono ancora state tradotte e vengono mostrate in inglese.
The Problem with Unstructured LLM Output
By default, LLMs return free-form text. Parsing that text to extract structured data is fragile: a change in model behavior, a slight prompt variation, or an edge case in the input can change the output format unexpectedly, breaking your parser and crashing your application.
Consider asking an LLM to 'return the user's name and age as JSON'. Sometimes it returns {"name":"Alice","age":30}, sometimes it wraps it in a markdown code block, sometimes it adds explanatory prose. Any of these variations requires different parsing logic. Reliable machine-readable output requires forcing the model to follow a structure, not hoping it does.
OpenAI JSON Mode
OpenAI introduced JSON mode via the response_format parameter. When set to {"type": "json_object"}, the model is constrained to always return a valid JSON object. The model will never output anything that is not valid JSON — no markdown wrappers, no explanatory text, no trailing prose.
import openai
import json
client = openai.OpenAI()
response = client.chat.completions.create(
model='gpt-4o-mini',
messages=[
{
'role': 'system',
'content': 'Extract information from the text and return valid JSON only.'
},
{
'role': 'user',
'content': 'John Smith, age 34, works as a software engineer in Austin.'
}
],
response_format={'type': 'json_object'} # Guarantee valid JSON output
)
# Safe to parse - guaranteed valid JSON
data = json.loads(response.choices[0].message.content)
print(data)
# Example output: {"name": "John Smith", "age": 34, "job": "software engineer", "city": "Austin"}JSON Mode Caveats
JSON mode guarantees valid JSON syntax but does NOT guarantee that the JSON contains the fields you want. The model still decides what keys to include, their names, and the data types it uses. You might ask for a name field and get back full_name instead, or ask for an array and get a string.
Also note: JSON mode requires that you mention JSON in your prompt. If you enable JSON mode but your prompt does not ask for JSON output, the model may produce an empty JSON object or refuse to generate. Always explicitly instruct the model to respond in JSON format in the system or user message.
Structured Outputs with Pydantic (Preview)
OpenAI's newer Structured Outputs feature goes further than JSON mode: you provide a JSON Schema, and the model is constrained to return exactly that schema — specific field names, types, and nesting. This eliminates the schema inconsistency problem of basic JSON mode.
The Python SDK accepts Pydantic models directly, automatically converting them to JSON Schema and deserializing the response back into a typed Python object. This is the cleanest way to get reliable structured data from an LLM in Python.
import openai
from pydantic import BaseModel
from typing import Optional
client = openai.OpenAI()
class PersonInfo(BaseModel):
name: str
age: Optional[int]
job_title: str
city: str
completion = client.beta.chat.completions.parse(
model='gpt-4o-mini',
messages=[
{'role': 'system', 'content': 'Extract person information from the text.'},
{'role': 'user', 'content': 'Sarah Chen, 28 years old, is a data scientist based in Seattle.'}
],
response_format=PersonInfo # Pass Pydantic model directly
)
# Already deserialized into a PersonInfo instance
person = completion.choices[0].message.parsed
print(person.name) # Sarah Chen
print(person.age) # 28
print(person.job_title) # data scientist
print(person.city) # SeattleCrafting Prompts for Consistent JSON
Even with JSON mode enabled, your prompt design affects output quality. Best practices for JSON prompts:
- Name the fields explicitly: Tell the model exactly what fields you expect, not just 'return JSON'
- Specify types: 'Return the price as a number, not a string' prevents type mismatches
- Define enumerations: 'The category must be one of: bug, feature, question' prevents unexpected values
- Handle missing data: 'If a field is not present in the text, return null for that field'
Think of your prompt as a partial JSON Schema written in prose. The more precisely you specify the output contract, the more reliably the model will follow it.
Reliable JSON Without Structured Outputs
If you are using a model that does not support structured outputs or JSON mode, you can still get reliable JSON by being very explicit in your prompt and parsing defensively. The key technique is to ask the model to wrap its JSON in XML tags, which makes extraction unambiguous regardless of any surrounding text.
import re
import json
import openai
client = openai.OpenAI()
def extract_json_from_response(text):
# Try direct parse first
try:
return json.loads(text)
except json.JSONDecodeError:
pass
# Try extracting from XML tags
match = re.search(r'<json>(.*?)</json>', text, re.DOTALL)
if match:
return json.loads(match.group(1))
# Try extracting from JSON object pattern
match = re.search(r'({.*})', text, re.DOTALL)
if match:
return json.loads(match.group(1))
raise ValueError('No valid JSON found in response')
prompt = ('Extract the product info as JSON with fields: name, price_usd, in_stock.\n'
'Wrap your JSON in <json></json> tags.\n\n'
'Product: Blue Wireless Headphones cost $89.99, in stock.')
resp = client.chat.completions.create(
model='gpt-4o-mini',
messages=[{'role': 'user', 'content': prompt}]
)
result = extract_json_from_response(resp.choices[0].message.content)
print(result)Nested JSON Structures
JSON mode and structured outputs handle arbitrarily nested structures. You can define Pydantic models with lists, nested objects, and optional fields, and the model will populate the full structure correctly.
from pydantic import BaseModel
from typing import List, Optional
import openai
client = openai.OpenAI()
class LineItem(BaseModel):
product: str
quantity: int
unit_price: float
class Invoice(BaseModel):
vendor: str
invoice_number: Optional[str]
line_items: List[LineItem]
total: float
raw_text = '''
INVOICE #INV-2025-0042
From: TechSupplies Inc.
- 3x USB Hubs at $24.99 each
- 1x 4K Monitor at $399.00
Total: $474.97
'''
completion = client.beta.chat.completions.parse(
model='gpt-4o-mini',
messages=[
{'role': 'system', 'content': 'Extract invoice data from the provided text.'},
{'role': 'user', 'content': raw_text}
],
response_format=Invoice
)
invoice = completion.choices[0].message.parsed
print(f'Vendor: {invoice.vendor}')
print(f'Items: {len(invoice.line_items)}')
print(f'Total: ${invoice.total}')Handling Refusals in Structured Mode
When using structured outputs, the model may sometimes refuse to complete the extraction — for example, if the input text is empty, harmful, or clearly does not contain the requested information. In structured outputs mode, refusals are indicated by the refusal field on the message rather than the parsed field.
Always check for refusals before accessing the parsed result, especially when processing user-provided or untrusted input that might trigger content filters.
import openai
from pydantic import BaseModel
client = openai.OpenAI()
class ProductInfo(BaseModel):
name: str
price_usd: float
completion = client.beta.chat.completions.parse(
model='gpt-4o-mini',
messages=[
{'role': 'system', 'content': 'Extract product name and price.'},
{'role': 'user', 'content': 'Tell me how to build a weapon.'}
],
response_format=ProductInfo
)
message = completion.choices[0].message
if message.refusal:
print('Model refused:', message.refusal)
else:
product = message.parsed
print(f'Name: {product.name}, Price: {product.price_usd}')JSON for Multi-Value Extraction
JSON mode is especially powerful for extracting multiple distinct pieces of information from a single piece of text in one API call, rather than making separate calls for each field. Extract all the fields you need at once and parse the result into your data model.
This reduces both API calls and cost compared to asking for one field at a time. A single well-structured extraction prompt can pull names, dates, monetary amounts, sentiment, action items, and classification labels all at once from a single document.
Streaming with JSON Mode
JSON mode is compatible with streaming, but with an important constraint: the JSON is only valid once the complete response has been streamed. Individual token chunks of JSON are not valid JSON on their own. This means you must accumulate the full streaming response before parsing when using JSON mode.
For streaming applications that also need JSON output, use structured outputs with streaming, accumulate all chunks, then parse when the stream finishes. Alternatively, design your streaming UI to show a loading state while the JSON accumulates, then render the parsed result.
When to Use JSON Mode vs Structured Outputs
Choose the right tool for your scenario:
- JSON mode: Simple cases, prototyping, or when you only need valid JSON syntax without strict field enforcement. Use when the model deciding field names is acceptable.
- Structured outputs with Pydantic: Production systems that parse results programmatically. Use when you need guaranteed field names, types, and nested structure. This is the recommended approach for any extraction pipeline.
- XML tag extraction: Fallback for models that do not support JSON mode or when you need to extract JSON embedded in a longer response.
Quick Check
Test your understanding of AI Engineering concepts from this lesson.
Lesson Recap
In this lesson you learned: JSON mode via response_format guarantees valid JSON syntax but not specific field schemas, structured outputs with Pydantic models enforce exact field names and types using JSON Schema, and always check for refusals before accessing parsed results when processing untrusted input. Next up we explore defining Pydantic schemas in depth for typed extraction from complex documents.
Domande Frequenti
La lezione «Modalità JSON e response_format» è gratuita?
Sì — il testo completo di «Modalità JSON e response_format» è gratuito qui sul web. Per esercitarvi in modo interattivo (un editor di codice integrato e un tutor IA 24/7) e sbloccare il resto del corso AI Engineering Academy, passa a CoddyKit PRO. Il corso AI Engineering Academy include 4 lezioni in totale.
Cosa imparerò in «Modalità JSON e response_format»?
Abiliterà la modalità JSON nell'API OpenAI, creerà prompt che producano costantemente JSON valido e gestirà i casi in cui il modello riesca comunque a rompere il formato. Eserciti AI Engineering Academy con codice pratico che esegui direttamente nel browser, e un tutor IA 24/7 risponde alle tue domande mentre lavori sulla lezione.
Ho bisogno di esperienza per iniziare AI Engineering Academy?
Non è richiesta alcuna esperienza precedente. AI Engineering Academy su CoddyKit è strutturato per principianti e studenti avanzati, quindi puoi iniziare da qui o dall'inizio e procedere al tuo ritmo. Questa è la lezione 1 di 4.
Quanto tempo richiede la lezione «Modalità JSON e response_format»?
La maggior parte delle lezioni CoddyKit richiede circa 5–10 minuti. Ogni lezione è breve e interattiva, quindi fai progressi costanti e riprendi esattamente da dove hai lasciato su web e app.
Posso scrivere ed eseguire codice in questa lezione AI Engineering Academy?
Sì. Ogni lezione AI Engineering Academy include un editor di codice integrato, quindi scrivi ed esegui codice reale direttamente nel tuo browser e ricevi feedback istantaneo dall'IA — nessuna configurazione locale necessaria.
Tutte le lezioni di questo corso
- Modalità JSON e response_format
- Output strutturati con Pydantic
- Estrarre dati da testo non strutturato
- Validare e riprovare gli output errati