Prompty do OCR i analizy dokumentów
Wyodrębniaj tekst, tabele i strukturę z obrazów dokumentów.
Prompty do OCR i analizy dokumentów to bezpłatna lekcja AI Prompt Engineering na CoddyKit. To lekcja 4 z 4. Możesz przeczytać całą lekcję poniżej za darmo — a potem ćwiczyć ją interaktywnie w przeglądarce z wbudowanym edytorem kodu i tutorem AI dostępnym 24/7. To część ścieżki edukacyjnej AI Prompt Engineering, a Twój postęp synchronizuje się między webem a aplikacją CoddyKit. Kurs AI Prompt Engineering zawiera 4 lekcji w sumie.
LLM-y jako czytniki dokumentów
OCR (optyczne rozpoznawanie znaków) tradycyjnie wymagało specjalistycznego oprogramowania do wyodrębniania tekstu z obrazów. Wizyjne modele LLM mogą teraz odczytywać tekst z obrazów, a także rozumieć jego treść — nie tylko wyodrębniać znaki, lecz także analizować strukturę, tabele, pismo odręczne i kontekst.
Typowe zadania związane z analizą dokumentów:
- Wyodrębnianie tekstu ze skanowanych dokumentów
- Odczytywanie paragonów, faktur i formularzy
- Analizowanie tabel i wykresów
- Transkrypcja odręcznych notatek
- Odczytywanie drukowanych etykiet i znaków
Podstawowy prompt do wyodrębniania tekstu
W przypadku prostego wyodrębniania tekstu z obrazu dokumentu:
import anthropic, base64
client = anthropic.Anthropic(api_key='YOUR_API_KEY')
def extract_text(image_path, extraction_prompt):
with open(image_path, 'rb') as f:
img_b64 = base64.standard_b64encode(f.read()).decode('utf-8')
r = client.messages.create(
model='claude-opus-4-5', max_tokens=1000,
messages=[{'role': 'user', 'content': [
{'type': 'image', 'source': {'type': 'base64', 'media_type': 'image/jpeg', 'data': img_b64}},
{'type': 'text', 'text': extraction_prompt}
]}]
)
return r.content[0].text
# Basic extraction
basic_prompt = 'Extract all text from this document image exactly as it appears. Preserve line breaks.'
# Structure-preserving extraction
structured_prompt = 'Extract all text from this document image. Preserve: paragraph structure, line breaks, and any visible formatting. Do not add any text not present in the image.'
print('Text extraction functions defined.')Zachowanie struktury tabeli
Gdy dokument zawiera tabelę, wyodrębnienie jej jako płaskiego tekstu powoduje utratę struktury. Należy używać promptów, które wyraźnie wymagają zachowania formatu tabeli:
table_prompt = '''
Extract all text from this document image.
If the document contains any tables, preserve the table structure using markdown table format:
| Column 1 | Column 2 | Column 3 |
|----------|----------|----------|
| Value | Value | Value |
For any text outside tables, use plain text preserving paragraph structure.
Do not invent or infer any data not visible in the image.
'''
# For structured output:
table_json_prompt = '''
Extract the table from this image.
Return JSON:
{
"headers": ["column name"],
"rows": [["cell value", "cell value"]],
"caption": "table caption if present or null"
}
If a cell is empty or illegible, use null.
'''
print('Table extraction prompts defined.')Pozycje na paragonie
Paragony mają określoną strukturę — nagłówek (sprzedawca), pozycje i podsumowania. Prompt przeznaczony do paragonów niezawodnie wyodrębnia tę strukturę:
import json
receipt_prompt = '''
Extract all information from this receipt image.
Return JSON:
{
"merchant": {
"name": str,
"address": str or null,
"phone": str or null
},
"transaction": {
"date": "YYYY-MM-DD or as written",
"time": "HH:MM or as written or null",
"receipt_number": str or null,
"payment_method": str or null
},
"items": [
{"description": str, "quantity": number or null, "unit_price": number or null, "total": number}
],
"subtotal": number or null,
"tax": number or null,
"tip": number or null,
"total": number,
"currency": "3-letter ISO code"
}
For any field not visible, use null. For numbers, use numeric type (not string).
'''
def extract_receipt(image_path):
text = extract_text(image_path, receipt_prompt)
return json.loads(text)
print('Receipt extraction function defined.')Transkrypcja odręcznej notatki
Transkrypcja treści odręcznej wymaga promptów uwzględniających związane z nią trudności — nieczytelne słowa, przekreślony tekst i skróty:
handwriting_prompt = '''
Transcribe the handwritten text in this image as accurately as possible.
Handling rules:
- If a word is illegible, write [ILLEGIBLE]
- If a word is partially legible, write [PARTIAL: best_guess]
- If text is crossed out, include it with strikethrough notation: ~~crossed out text~~
- Preserve line breaks as they appear
- If there are arrows, circles, or annotations, note them in brackets: [arrow pointing right]
- Do not correct spelling or grammar
After transcription, estimate overall legibility: high (>90% readable) | medium (70-90%) | low (<70%)
Format:
TRANSCRIPTION:
[transcribed text here]
LEGIBILITY: [rating]
'''
print(handwriting_prompt)Wyodrębnianie pól formularza
Formularze drukowane zawierają opisane pola oraz wprowadzone wartości. Prompt do wyodrębniania danych z formularza mapuje etykiety na wartości:
form_prompt = '''
Extract all form fields and their values from this document image.
For each field:
- Field label: the printed label (e.g., "First Name:", "Date of Birth:")
- Field value: the filled-in value (handwritten or typed)
- Filled: whether the field has been filled in (true/false)
Return JSON:
{
"form_title": str or null,
"fields": [
{
"label": str,
"value": str or null,
"filled": true | false
}
],
"signature_present": true | false,
"date_signed": str or null
}
If the value is illegible, use "[ILLEGIBLE]".
If the field is blank, value should be null and filled should be false.
'''
print('Form field extraction prompt defined.')
print('Handles: printed forms, questionnaires, applications.')Klasyfikacja dokumentu przed wyodrębnianiem danych
Przed wyodrębnianiem danych należy dodać etap klasyfikacji, aby zastosować odpowiedni schemat dla każdego typu dokumentu:
import json
DOC_SCHEMAS = {
'receipt': receipt_prompt,
'form': form_prompt,
'table': table_json_prompt,
'letter': 'Extract all text preserving paragraph structure. Identify: sender, recipient, date, subject, body.',
'label': 'Extract all text from this label. Include: product name, ingredients/contents, weight, expiry date, barcode numbers.'
}
def classify_and_extract(image_path):
# Step 1: Classify document type
classify_prompt = 'What type of document is this? Return JSON: {"type": "receipt|form|table|letter|label|other", "confidence": "high|medium|low"}'
classification_text = extract_text(image_path, classify_prompt)
doc_type = json.loads(classification_text)['type']
# Step 2: Apply correct schema
schema = DOC_SCHEMAS.get(doc_type, 'Extract all visible text from this document.')
extracted = extract_text(image_path, schema)
return {'type': doc_type, 'data': extracted}
print('Document classify-then-extract pipeline defined.')Obsługa obrazów niskiej jakości
Nie wszystkie obrazy dokumentów są wyraźne. Prompty powinny zapewniać prawidłową obsługę obrazów o obniżonej jakości:
low_quality_prompt = '''
Extract text from this document image. The image may be low quality, blurry, or poorly lit.
Extraction guidelines:
- Extract all text you can read with reasonable confidence
- For unclear sections, use [UNCLEAR] as a placeholder
- For completely unreadable sections, use [UNREADABLE: approximately N words]
- Do not guess or hallucinate words you cannot see clearly
- Note image quality issues at the end: "Image quality: [good/fair/poor]. Issues: [description]"
Be conservative — it is better to mark something as unclear than to guess incorrectly.
'''
print(low_quality_prompt)
print('\nConservative approach: unclear beats hallucinated.')Podsumowanie dokumentu wielostronicowego
W przypadku wielostronicowych dokumentów przesyłanych jako wiele obrazów należy połączyć wyodrębnianie danych strona po stronie z etapem syntezy:
def extract_multi_page_document(image_paths):
# Step 1: Extract text from each page
page_texts = []
for i, path in enumerate(image_paths):
page_text = extract_text(path, f'Extract all text from page {i+1} of this document. Preserve structure.')
page_texts.append(f'=== PAGE {i+1} ===\n{page_text}')
full_text = '\n\n'.join(page_texts)
# Step 2: Synthesize summary and key information
r = client.messages.create(
model='claude-opus-4-5', max_tokens=500,
messages=[{'role': 'user', 'content': f'''
Here is the extracted text from a {len(image_paths)}-page document:\n\n{full_text}\n\n
Provide:
1. Document type and title
2. 3-sentence summary
3. Key data points extracted
Return JSON: {{"type": str, "title": str, "summary": str, "key_data": [str]}}
'''}]
)
return json.loads(r.content[0].text)
print('Multi-page document pipeline defined.')Weryfikacja po wyodrębnieniu danych przez OCR
Przed wykorzystaniem wyników OCR w systemach zależnych należy je zweryfikować. Typowe kontrole poprawności:
import re
from datetime import datetime
def validate_receipt_extraction(data):
errors = []
# Validate total is present and numeric
if data.get('total') is None:
errors.append('total is missing')
elif not isinstance(data['total'], (int, float)):
errors.append(f'total is not numeric: {data["total"]}')
# Validate date format
if data.get('transaction', {}).get('date'):
date_str = data['transaction']['date']
try:
datetime.strptime(date_str, '%Y-%m-%d')
except ValueError:
errors.append(f'date format invalid: {date_str}')
# Validate line items total approximately equals subtotal
if data.get('items') and data.get('subtotal'):
items_total = sum(item.get('total', 0) for item in data['items'] if item.get('total'))
if abs(items_total - data['subtotal']) > 0.05:
errors.append(f'Items total {items_total} does not match subtotal {data["subtotal"]}')
return errors
print('Receipt validation function defined.')Wyodrębnianie danych strukturalnych z wykresów
Wykresy i diagramy na obrazach dokumentów zawierają dane niewidoczne dla tradycyjnego OCR, ale możliwe do odczytania przez wizyjne modele LLM. Prompt do wyodrębniania danych z wykresu wymaga od modelu odczytania leżących u podstaw wartości danych:
chart_prompt = '''
Extract the data from this chart or graph image.
Identify:
1. Chart type (bar, line, pie, scatter, table)
2. Title and axis labels
3. All data series names
4. All data points with their labels/values
5. Any notable trend or pattern
Return JSON:
{
"chart_type": str,
"title": str or null,
"x_axis_label": str or null,
"y_axis_label": str or null,
"data_series": [
{"name": str, "values": [{"label": str, "value": number}]}
],
"key_insight": str
}
If exact values are not readable, provide best estimates with a note.
'''
print(chart_prompt)Szybkie sprawdzenie
Jakie podejście zaleca się stosować w przypadku nieczytelnych słów podczas transkrypcji odręcznej treści z obrazu?
OCR i analiza dokumentów — najważniejsze wnioski
Wizyjne modele LLM zapewniają elastyczną analizę dokumentów wykraczającą poza rozpoznawanie znaków:
- Podstawowe wyodrębnianie: należy zachować podziały wierszy i strukturę akapitów oraz wyraźnie to określić
- Tabele: należy używać formatu tabeli Markdown albo schematu JSON z wierszami i nagłówkami, aby zachować strukturę
- Paragony: należy używać dedykowanego schematu z polami sprzedawcy, pozycji i podsumowań
- Pismo odręczne: należy używać oznaczenia [ILLEGIBLE] dla treści nieczytelnych i [PARTIAL] dla częściowo czytelnych — nigdy nie zgadywać
- Formularze: należy mapować pary etykieta–wartość oraz uwzględniać dla każdego pola status wypełnienia
- Najpierw należy sklasyfikować typ dokumentu, a następnie zastosować odpowiedni schemat wyodrębniania danych
- Wyodrębnione dane należy programowo weryfikować: sprawdzać wymagane pola, typy liczbowe, formaty dat i poprawność obliczeń
- Obrazy niskiej jakości: zachowawcze wyodrębnianie danych jest lepsze niż pewna siebie halucynacja
Ucz się AI Prompt Engineering dzięki korepetycjom AI — za darmo
Pisz i uruchamiaj kod w przeglądarce, otrzymuj natychmiastową pomoc od korepetytora AI dostępnego 24/7 i kontynuuj naukę w sieci lub w aplikacji.
- Kursy
- 53
- Lekcje
- 199
Często zadawane pytania
Czy lekcja „Prompty do OCR i analizy dokumentów” jest bezpłatna?
Tak — pełny tekst „Prompty do OCR i analizy dokumentów” jest dostępny za darmo tutaj w sieci. Aby ćwiczyć ją interaktywnie (wbudowany edytor kodu i tutor AI dostępny 24/7) i odblokować resztę kursu AI Prompt Engineering, przejdź na CoddyKit PRO. Kurs AI Prompt Engineering zawiera 4 lekcji w sumie.
Co nauczysz się w „Prompty do OCR i analizy dokumentów”?
Wyodrębniaj tekst, tabele i strukturę z obrazów dokumentów. Ćwiczysz AI Prompt Engineering z praktycznym kodem, który uruchamiasz bezpośrednio w przeglądarce, a tutor AI dostępny 24/7 odpowiada na Twoje pytania podczas pracy nad lekcją.
Czy potrzebuję doświadczenia, aby zacząć AI Prompt Engineering?
Nie wymagamy żadnego doświadczenia. AI Prompt Engineering w CoddyKit jest strukturyzowany dla początkujących i zaawansowanych użytkowników, więc możesz zacząć tutaj lub od początku i uczyć się w swoim tempie. To lekcja 4 z 4.
Ile czasu zajmuje lekcja „Prompty do OCR i analizy dokumentów”?
Większość lekcji CoddyKit trwa około 5–10 minut. Każda lekcja to mały, interaktywny krok, dzięki czemu robisz systematyczne postępy i zawsze wracasz dokładnie do tego samego miejsca — na webie i w aplikacji.
Czy mogę pisać i uruchamiać kod w tej lekcji AI Prompt Engineering?
Tak. Każda lekcja AI Prompt Engineering zawiera wbudowany edytor kodu, więc piszesz i uruchamiasz prawdziwy kod bezpośrednio w przeglądarce i od razu otrzymujesz sprzężenie zwrotne od AI — bez konfiguracji na komputerze.
Wszystkie lekcje w tym kursie
- Prompty do opisu obrazów i tworzenia podpisów
- Wizualne odpowiadanie na pytania
- Prompty do porównywania wielu obrazów
- Prompty do OCR i analizy dokumentów