Parsowanie i walidacja ustrukturyzowanych danych wyjściowych
Wymuszaj na LLM-ach zwracanie wiarygodnych danych strukturalnych za pomocą parserów outputu i schematów, a gdy model wygeneruje nieprawidłowy output, przeprowadzaj walidację lub ponawiaj próbę.
Parsowanie i walidacja ustrukturyzowanych danych wyjściowych to bezpłatna lekcja AI Agents with LangChain & Autonomous Workflows na CoddyKit. To lekcja 4 z 4. Możesz przeczytać całą lekcję poniżej za darmo — a potem ćwiczyć ją interaktywnie w przeglądarce z wbudowanym edytorem kodu i tutorem AI dostępnym 24/7. To część ścieżki edukacyjnej AI Agents with LangChain & Autonomous Workflows, a Twój postęp synchronizuje się między webem a aplikacją CoddyKit. Kurs AI Agents with LangChain & Autonomous Workflows zawiera 4 lekcji w sumie.
Części tej lekcji nie zostały jeszcze przetłumaczone i są wyświetlane po angielsku.
The Problem with Free Text
LLMs return prose by default, but your code needs structured data: JSON, a list, a typed object. Parsing free text with regex is fragile.
This lesson covers getting reliable structured output from models.
Asking for a Format
The first step is simply instructing the model to produce a specific format. But instruction alone is not enough; models drift, add prose, or wrap output in markdown.
prompt = 'Extract name and age as JSON: "Lena is 30"'
# model might reply: 'Sure! {"name":"Lena","age":30}'Output Parsers
LangChain output parsers do two jobs: they generate format instructions to inject into the prompt, and they parse the model's response back into a structured object.
from langchain.output_parsers import CommaSeparatedListOutputParser
parser = CommaSeparatedListOutputParser()
print(parser.get_format_instructions())Schema-Based Parsing
Define the shape you want with a schema (e.g. a Pydantic model). The parser turns it into instructions and validates the result against the fields and types.
from pydantic import BaseModel
class Person(BaseModel):
name: str
age: intInjecting Format Instructions
Add the parser's instructions into your prompt template so the model knows exactly what structure to emit.
template = 'Extract info.\n{format_instructions}\nText: {text}'
prompt = template.format(
format_instructions=parser.get_format_instructions(),
text='Lena is 30')Parsing the Response
After the model replies, the parser converts the text into your typed object, raising an error if it does not match the schema.
result = parser.parse(model_output)
print(result.name, result.age)Handling Malformed Output
Models occasionally produce invalid JSON. A retry/fixing parser detects the failure and asks the model to correct its own output, turning a hard crash into a recoverable step.
from langchain.output_parsers import RetryOutputParser
robust = RetryOutputParser.from_llm(parser=parser, llm=llm)Native JSON / Tool Modes
Many modern models support a JSON mode or function/tool calling that constrains output to valid structured data at the API level. When available, this is far more reliable than prompt instructions alone.
Validation Beyond Types
A value can be the right type but still wrong: a negative age, an empty required field. Add validators so business rules are enforced, not just the data shape.
if result.age < 0 or result.age > 130:
raise ValueError('age out of range')Why It Matters for Agents
Agents chain steps together, feeding one output into the next. If a step emits malformed data, the whole chain breaks. Structured, validated output is what makes multi-step agents dependable.
A Reliable Output Workflow
Putting it together:
- Define a schema for the data you need
- Inject format instructions into the prompt
- Prefer native JSON/tool mode when available
- Parse and validate, with a retry parser as a safety net
Quick Check
Test your understanding of structured output.
Recap
You learned to get reliable structured data from LLMs.
- Output parsers generate instructions and parse responses
- Schemas validate shape and types
- Retry parsers recover from malformed output
- Native JSON/tool modes are most reliable when available
Często zadawane pytania
Czy lekcja „Parsowanie i walidacja ustrukturyzowanych danych wyjściowych” jest bezpłatna?
Tak — pełny tekst „Parsowanie i walidacja ustrukturyzowanych danych wyjściowych” jest dostępny za darmo tutaj w sieci. Aby ćwiczyć ją interaktywnie (wbudowany edytor kodu i tutor AI dostępny 24/7) i odblokować resztę kursu AI Agents with LangChain & Autonomous Workflows, przejdź na CoddyKit PRO. Kurs AI Agents with LangChain & Autonomous Workflows zawiera 4 lekcji w sumie.
Co nauczysz się w „Parsowanie i walidacja ustrukturyzowanych danych wyjściowych”?
Wymuszaj na LLM-ach zwracanie wiarygodnych danych strukturalnych za pomocą parserów outputu i schematów, a gdy model wygeneruje nieprawidłowy output, przeprowadzaj walidację lub ponawiaj próbę. Ćwiczysz AI Agents with LangChain & Autonomous Workflows z praktycznym kodem, który uruchamiasz bezpośrednio w przeglądarce, a tutor AI dostępny 24/7 odpowiada na Twoje pytania podczas pracy nad lekcją.
Czy potrzebuję doświadczenia, aby zacząć AI Agents with LangChain & Autonomous Workflows?
Nie wymagamy żadnego doświadczenia. AI Agents with LangChain & Autonomous Workflows w CoddyKit jest strukturyzowany dla początkujących i zaawansowanych użytkowników, więc możesz zacząć tutaj lub od początku i uczyć się w swoim tempie. To lekcja 4 z 4.
Ile czasu zajmuje lekcja „Parsowanie i walidacja ustrukturyzowanych danych wyjściowych”?
Większość lekcji CoddyKit trwa około 5–10 minut. Każda lekcja to mały, interaktywny krok, dzięki czemu robisz systematyczne postępy i zawsze wracasz dokładnie do tego samego miejsca — na webie i w aplikacji.
Czy mogę pisać i uruchamiać kod w tej lekcji AI Agents with LangChain & Autonomous Workflows?
Tak. Każda lekcja AI Agents with LangChain & Autonomous Workflows zawiera wbudowany edytor kodu, więc piszesz i uruchamiasz prawdziwy kod bezpośrednio w przeglądarce i od razu otrzymujesz sprzężenie zwrotne od AI — bez konfiguracji na komputerze.
Wszystkie lekcje w tym kursie
- Skuteczne techniki projektowania promptów
- Integracja LLM-ów z LangChain
- Zarządzanie parametrami modeli i kosztami
- Parsowanie i walidacja ustrukturyzowanych danych wyjściowych