AI-promptengineering · Les

Prompts voor OCR en documentanalyse

Haal tekst, tabellen en structuur uit afbeeldingen van documenten.

Les 4 van 413 stappen

Prompts voor OCR en documentanalyse is een gratis AI-promptengineering-les op CoddyKit. Dit is les 4 van 4. Je kunt 3 lessen uit dit leerpad gratis volledig lezen — daarna ontgrendelt CoddyKit PRO alle lessen, plus praktische oefeningen met een ingebouwde code-editor en een AI-tutor die 24/7 beschikbaar is. Deze les maakt deel uit van het leertraject AI-promptengineering. Je voortgang wordt gesynchroniseerd op het web en in de CoddyKit-app. De cursus AI-promptengineering bevat in totaal 4 lessen.

LLM's als documentlezers

OCR (Optische tekenherkenning) vereiste traditioneel gespecialiseerde software om tekst uit afbeeldingen te halen. Vision-LLM's kunnen nu tekst uit afbeeldingen lezen en de inhoud ook begrijpen — ze halen niet alleen tekens eruit, maar interpreteren ook structuur, tabellen, handschrift en context.

Veelvoorkomende taken voor documentanalyse:

  • Tekst uit gescande documenten halen
  • Bonnen, facturen en formulieren lezen
  • Tabellen en grafieken ontleden
  • Handgeschreven notities transcriberen
  • Gedrukte labels en borden lezen

Basisprompt voor tekstextractie

Voor eenvoudige tekstextractie uit een documentafbeelding:

import anthropic, base64

client = anthropic.Anthropic(api_key='YOUR_API_KEY')

def extract_text(image_path, extraction_prompt):
    with open(image_path, 'rb') as f:
        img_b64 = base64.standard_b64encode(f.read()).decode('utf-8')

    r = client.messages.create(
        model='claude-opus-4-5', max_tokens=1000,
        messages=[{'role': 'user', 'content': [
            {'type': 'image', 'source': {'type': 'base64', 'media_type': 'image/jpeg', 'data': img_b64}},
            {'type': 'text', 'text': extraction_prompt}
        ]}]
    )
    return r.content[0].text

# Basic extraction
basic_prompt = 'Extract all text from this document image exactly as it appears. Preserve line breaks.'

# Structure-preserving extraction
structured_prompt = 'Extract all text from this document image. Preserve: paragraph structure, line breaks, and any visible formatting. Do not add any text not present in the image.'

print('Text extraction functions defined.')

Tabelstructuur behouden

Wanneer een document een tabel bevat, gaat de structuur verloren als je deze als platte tekst extraheert. Gebruik prompts die de tabelindeling expliciet behouden:

table_prompt = '''
Extract all text from this document image.
If the document contains any tables, preserve the table structure using markdown table format:
| Column 1 | Column 2 | Column 3 |
|----------|----------|----------|
| Value    | Value    | Value    |

For any text outside tables, use plain text preserving paragraph structure.
Do not invent or infer any data not visible in the image.
'''

# For structured output:
table_json_prompt = '''
Extract the table from this image.
Return JSON:
{
  "headers": ["column name"],
  "rows": [["cell value", "cell value"]],
  "caption": "table caption if present or null"
}
If a cell is empty or illegible, use null.
'''

print('Table extraction prompts defined.')

Bonartikelen uitsplitsen

Bonnen hebben een specifieke structuur — kop (verkoper), regelitems en totalen. Een bon-specifieke prompt haalt deze structuur betrouwbaar uit:

import json

receipt_prompt = '''
Extract all information from this receipt image.
Return JSON:
{
  "merchant": {
    "name": str,
    "address": str or null,
    "phone": str or null
  },
  "transaction": {
    "date": "YYYY-MM-DD or as written",
    "time": "HH:MM or as written or null",
    "receipt_number": str or null,
    "payment_method": str or null
  },
  "items": [
    {"description": str, "quantity": number or null, "unit_price": number or null, "total": number}
  ],
  "subtotal": number or null,
  "tax": number or null,
  "tip": number or null,
  "total": number,
  "currency": "3-letter ISO code"
}
For any field not visible, use null. For numbers, use numeric type (not string).
'''

def extract_receipt(image_path):
    text = extract_text(image_path, receipt_prompt)
    return json.loads(text)

print('Receipt extraction function defined.')

Handgeschreven notities transcriberen

Voor het transcriberen van handgeschreven inhoud zijn prompts nodig die rekening houden met de uitdagingen — onleesbare woorden, doorgehaalde tekst en afkortingen:

handwriting_prompt = '''
Transcribe the handwritten text in this image as accurately as possible.

Handling rules:
- If a word is illegible, write [ILLEGIBLE]
- If a word is partially legible, write [PARTIAL: best_guess]
- If text is crossed out, include it with strikethrough notation: ~~crossed out text~~
- Preserve line breaks as they appear
- If there are arrows, circles, or annotations, note them in brackets: [arrow pointing right]
- Do not correct spelling or grammar

After transcription, estimate overall legibility: high (>90% readable) | medium (70-90%) | low (<70%)

Format:
TRANSCRIPTION:
[transcribed text here]

LEGIBILITY: [rating]
'''

print(handwriting_prompt)

Formuliervelden extraheren

Gedrukte formulieren hebben gelabelde velden en ingevulde waarden. Een prompt voor formulierextractie koppelt labels aan waarden:

form_prompt = '''
Extract all form fields and their values from this document image.

For each field:
- Field label: the printed label (e.g., "First Name:", "Date of Birth:")
- Field value: the filled-in value (handwritten or typed)
- Filled: whether the field has been filled in (true/false)

Return JSON:
{
  "form_title": str or null,
  "fields": [
    {
      "label": str,
      "value": str or null,
      "filled": true | false
    }
  ],
  "signature_present": true | false,
  "date_signed": str or null
}

If the value is illegible, use "[ILLEGIBLE]".
If the field is blank, value should be null and filled should be false.
'''

print('Form field extraction prompt defined.')
print('Handles: printed forms, questionnaires, applications.')

Documenten classificeren vóór extractie

Voeg vóór de extractie een classificatiestap toe, zodat voor elk documenttype het juiste extractieschema kan worden toegepast:

import json

DOC_SCHEMAS = {
    'receipt': receipt_prompt,
    'form': form_prompt,
    'table': table_json_prompt,
    'letter': 'Extract all text preserving paragraph structure. Identify: sender, recipient, date, subject, body.',
    'label': 'Extract all text from this label. Include: product name, ingredients/contents, weight, expiry date, barcode numbers.'
}

def classify_and_extract(image_path):
    # Step 1: Classify document type
    classify_prompt = 'What type of document is this? Return JSON: {"type": "receipt|form|table|letter|label|other", "confidence": "high|medium|low"}'
    classification_text = extract_text(image_path, classify_prompt)
    doc_type = json.loads(classification_text)['type']

    # Step 2: Apply correct schema
    schema = DOC_SCHEMAS.get(doc_type, 'Extract all visible text from this document.')
    extracted = extract_text(image_path, schema)

    return {'type': doc_type, 'data': extracted}

print('Document classify-then-extract pipeline defined.')

Omgaan met afbeeldingen van lage kwaliteit

Niet alle documentafbeeldingen zijn duidelijk. Prompts moeten op gepaste wijze omgaan met verslechterde afbeeldingen:

low_quality_prompt = '''
Extract text from this document image. The image may be low quality, blurry, or poorly lit.

Extraction guidelines:
- Extract all text you can read with reasonable confidence
- For unclear sections, use [UNCLEAR] as a placeholder
- For completely unreadable sections, use [UNREADABLE: approximately N words]
- Do not guess or hallucinate words you cannot see clearly
- Note image quality issues at the end: "Image quality: [good/fair/poor]. Issues: [description]"

Be conservative — it is better to mark something as unclear than to guess incorrectly.
'''

print(low_quality_prompt)
print('\nConservative approach: unclear beats hallucinated.')

Samenvatting van een document met meerdere pagina's

Combineer bij documenten met meerdere pagina's die als meerdere afbeeldingen worden verzonden de extractie per pagina met een synthesestap:

def extract_multi_page_document(image_paths):
    # Step 1: Extract text from each page
    page_texts = []
    for i, path in enumerate(image_paths):
        page_text = extract_text(path, f'Extract all text from page {i+1} of this document. Preserve structure.')
        page_texts.append(f'=== PAGE {i+1} ===\n{page_text}')

    full_text = '\n\n'.join(page_texts)

    # Step 2: Synthesize summary and key information
    r = client.messages.create(
        model='claude-opus-4-5', max_tokens=500,
        messages=[{'role': 'user', 'content': f'''
Here is the extracted text from a {len(image_paths)}-page document:\n\n{full_text}\n\n
Provide:
1. Document type and title
2. 3-sentence summary
3. Key data points extracted
Return JSON: {{"type": str, "title": str, "summary": str, "key_data": [str]}}
'''}]
    )
    return json.loads(r.content[0].text)

print('Multi-page document pipeline defined.')

Validatie na OCR-extractie

OCR-uitvoer moet worden gevalideerd voordat deze in vervolgsystemen wordt gebruikt. Veelvoorkomende validatiecontroles:

import re
from datetime import datetime

def validate_receipt_extraction(data):
    errors = []

    # Validate total is present and numeric
    if data.get('total') is None:
        errors.append('total is missing')
    elif not isinstance(data['total'], (int, float)):
        errors.append(f'total is not numeric: {data["total"]}')

    # Validate date format
    if data.get('transaction', {}).get('date'):
        date_str = data['transaction']['date']
        try:
            datetime.strptime(date_str, '%Y-%m-%d')
        except ValueError:
            errors.append(f'date format invalid: {date_str}')

    # Validate line items total approximately equals subtotal
    if data.get('items') and data.get('subtotal'):
        items_total = sum(item.get('total', 0) for item in data['items'] if item.get('total'))
        if abs(items_total - data['subtotal']) > 0.05:
            errors.append(f'Items total {items_total} does not match subtotal {data["subtotal"]}')

    return errors

print('Receipt validation function defined.')

Gestructureerde gegevens uit diagrammen en grafieken halen

Diagrammen en grafieken in documentafbeeldingen bevatten gegevens die voor traditionele OCR onzichtbaar maar voor vision-LLM's leesbaar zijn. Een prompt voor grafiekextractie vraagt het model om de onderliggende gegevenswaarden te lezen:

chart_prompt = '''
Extract the data from this chart or graph image.

Identify:
1. Chart type (bar, line, pie, scatter, table)
2. Title and axis labels
3. All data series names
4. All data points with their labels/values
5. Any notable trend or pattern

Return JSON:
{
  "chart_type": str,
  "title": str or null,
  "x_axis_label": str or null,
  "y_axis_label": str or null,
  "data_series": [
    {"name": str, "values": [{"label": str, "value": number}]}
  ],
  "key_insight": str
}

If exact values are not readable, provide best estimates with a note.
'''

print(chart_prompt)

Korte controle

Wat is bij het transcriberen van handgeschreven inhoud uit een afbeelding de aanbevolen aanpak voor onleesbare woorden?

OCR en documentanalyse — belangrijkste inzichten

Vision-LLM's bieden flexibele documentanalyse die verder gaat dan tekenherkenning:

  • Basisextractie: behoud regeleinden en alinea-structuur; geef dit expliciet op
  • Tabellen: gebruik een markdown-tabelindeling of een schema voor JSON-rijen en -kopteksten om de structuur te behouden
  • Bonnen: gebruik een speciaal schema met velden voor verkoper, artikelen en totalen
  • Handschrift: gebruik [ONLEESBAAR] voor onleesbare tekst en [GEDEELTELIJK] voor gedeeltelijk leesbare tekst — raad nooit
  • Formulieren: koppel label-waardeparen; neem voor elk veld de status ingevuld/niet ingevuld op
  • Classificeer eerst het documenttype en pas daarna het juiste extractieschema toe
  • Valideer geëxtraheerde gegevens programmatisch: verplichte velden, numerieke typen, datumindelingen en rekencontroles
  • Afbeeldingen van lage kwaliteit: een behoudende extractie is beter dan een zelfverzekerde hallucinatie
Gratis beginnen

Leer AI-promptengineering met een AI-tutor — gratis

Schrijf echte code en voer die uit in je browser, krijg direct hulp van een AI-tutor die 24/7 beschikbaar is en ga verder waar je gebleven bent op het web of in de app.

Cursussen
53
Lessen
199

Veelgestelde vragen

Is de les “Prompts voor OCR en documentanalyse” gratis?

Ja — je kunt hier op het web alle 3 lessen van het leerpad AI-promptengineering, waaronder “Prompts voor OCR en documentanalyse”, gratis volledig lezen. Daarna ontgrendelt CoddyKit PRO alle lessen, plus interactieve oefeningen met een ingebouwde code-editor en een AI-tutor die 24/7 beschikbaar is. De cursus AI-promptengineering bevat in totaal 4 lessen.

Wat leer ik in “Prompts voor OCR en documentanalyse”?

Haal tekst, tabellen en structuur uit afbeeldingen van documenten. Je oefent met AI-promptengineering door code rechtstreeks in de browser uit te voeren. Een AI-begeleider die 24/7 beschikbaar is beantwoordt je vragen terwijl je de les doorwerkt.

Heb ik ervaring nodig om met AI-promptengineering te beginnen?

Ervaring vooraf is niet nodig. AI-promptengineering op CoddyKit is opgebouwd voor beginners tot gevorderden, zodat je hier of bij het begin kunt starten en in je eigen tempo kunt leren. Dit is les 4 van 4.

Hoe lang duurt de les “Prompts voor OCR en documentanalyse”?

De meeste lessen van CoddyKit duren ongeveer 5–10 minuten. Elke les is kort en interactief, zodat je gestaag vooruitgaat en op het web en in de app precies verdergaat waar je was gebleven.

Kan ik code schrijven en uitvoeren in deze les over AI-promptengineering?

Ja. Elke les over AI-promptengineering bevat een ingebouwde code-editor, zodat je rechtstreeks in je browser echte code kunt schrijven en uitvoeren en direct feedback van AI krijgt — lokale installatie is niet nodig.

Alle lessen in deze cursus

  1. Prompts voor beeldbeschrijving en bijschriften
  2. Visuele vraagbeantwoording
  3. Prompts voor vergelijking van meerdere afbeeldingen
  4. Prompts voor OCR en documentanalyse
← Terug naar AI-promptengineering