AI Prompt Engineering · Lektion

Prompts für OCR und Dokumentenanalyse

Extrahieren Sie Text, Tabellen und Strukturen aus Dokumentbildern

Lektion 4 von 413 Schritte

Prompts für OCR und Dokumentenanalyse ist eine kostenlose AI Prompt Engineering-Lektion auf CoddyKit. Dies ist Lektion 4 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des AI Prompt Engineering-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der AI Prompt Engineering-Kurs umfasst insgesamt 4 Lektionen.

LLMs als Dokumentleser

OCR (optische Zeichenerkennung) erforderte traditionell spezialisierte Software, um Text aus Bildern zu extrahieren. Vision-LLMs können jetzt Text aus Bildern lesen und den Inhalt außerdem verstehen – nicht nur Zeichen extrahieren, sondern auch Struktur, Tabellen, Handschrift und Kontext erfassen.

Häufige Aufgaben der Dokumentanalyse:

  • Text aus gescannten Dokumenten extrahieren
  • Belege, Rechnungen und Formulare lesen
  • Tabellen und Diagramme verarbeiten
  • Handschriftliche Notizen transkribieren
  • Gedruckte Beschriftungen und Schilder lesen

Prompt zur grundlegenden Textextraktion

Für die einfache Textextraktion aus einem Dokumentbild:

import anthropic, base64

client = anthropic.Anthropic(api_key='YOUR_API_KEY')

def extract_text(image_path, extraction_prompt):
    with open(image_path, 'rb') as f:
        img_b64 = base64.standard_b64encode(f.read()).decode('utf-8')

    r = client.messages.create(
        model='claude-opus-4-5', max_tokens=1000,
        messages=[{'role': 'user', 'content': [
            {'type': 'image', 'source': {'type': 'base64', 'media_type': 'image/jpeg', 'data': img_b64}},
            {'type': 'text', 'text': extraction_prompt}
        ]}]
    )
    return r.content[0].text

# Basic extraction
basic_prompt = 'Extract all text from this document image exactly as it appears. Preserve line breaks.'

# Structure-preserving extraction
structured_prompt = 'Extract all text from this document image. Preserve: paragraph structure, line breaks, and any visible formatting. Do not add any text not present in the image.'

print('Text extraction functions defined.')

Erhalt der Tabellenstruktur

Wenn ein Dokument eine Tabelle enthält, geht beim Extrahieren als Fließtext die Struktur verloren. Verwenden Sie Prompts, die ausdrücklich den Erhalt des Tabellenformats verlangen:

table_prompt = '''
Extract all text from this document image.
If the document contains any tables, preserve the table structure using markdown table format:
| Column 1 | Column 2 | Column 3 |
|----------|----------|----------|
| Value    | Value    | Value    |

For any text outside tables, use plain text preserving paragraph structure.
Do not invent or infer any data not visible in the image.
'''

# For structured output:
table_json_prompt = '''
Extract the table from this image.
Return JSON:
{
  "headers": ["column name"],
  "rows": [["cell value", "cell value"]],
  "caption": "table caption if present or null"
}
If a cell is empty or illegible, use null.
'''

print('Table extraction prompts defined.')

Aufschlüsselung von Belegen

Belege weisen eine bestimmte Struktur auf – Kopfzeile (Händler), Einzelpositionen und Summen. Ein auf Belege zugeschnittener Prompt extrahiert diese Struktur zuverlässig:

import json

receipt_prompt = '''
Extract all information from this receipt image.
Return JSON:
{
  "merchant": {
    "name": str,
    "address": str or null,
    "phone": str or null
  },
  "transaction": {
    "date": "YYYY-MM-DD or as written",
    "time": "HH:MM or as written or null",
    "receipt_number": str or null,
    "payment_method": str or null
  },
  "items": [
    {"description": str, "quantity": number or null, "unit_price": number or null, "total": number}
  ],
  "subtotal": number or null,
  "tax": number or null,
  "tip": number or null,
  "total": number,
  "currency": "3-letter ISO code"
}
For any field not visible, use null. For numbers, use numeric type (not string).
'''

def extract_receipt(image_path):
    text = extract_text(image_path, receipt_prompt)
    return json.loads(text)

print('Receipt extraction function defined.')

Transkription handschriftlicher Notizen

Die Transkription handschriftlicher Inhalte erfordert Prompts, die die damit verbundenen Herausforderungen berücksichtigen – unleserliche Wörter, durchgestrichener Text und Abkürzungen:

handwriting_prompt = '''
Transcribe the handwritten text in this image as accurately as possible.

Handling rules:
- If a word is illegible, write [ILLEGIBLE]
- If a word is partially legible, write [PARTIAL: best_guess]
- If text is crossed out, include it with strikethrough notation: ~~crossed out text~~
- Preserve line breaks as they appear
- If there are arrows, circles, or annotations, note them in brackets: [arrow pointing right]
- Do not correct spelling or grammar

After transcription, estimate overall legibility: high (>90% readable) | medium (70-90%) | low (<70%)

Format:
TRANSCRIPTION:
[transcribed text here]

LEGIBILITY: [rating]
'''

print(handwriting_prompt)

Extraktion von Formularfeldern

Gedruckte Formulare enthalten beschriftete Felder und eingetragene Werte. Ein Prompt zur Formularextraktion ordnet Beschriftungen den Werten zu:

form_prompt = '''
Extract all form fields and their values from this document image.

For each field:
- Field label: the printed label (e.g., "First Name:", "Date of Birth:")
- Field value: the filled-in value (handwritten or typed)
- Filled: whether the field has been filled in (true/false)

Return JSON:
{
  "form_title": str or null,
  "fields": [
    {
      "label": str,
      "value": str or null,
      "filled": true | false
    }
  ],
  "signature_present": true | false,
  "date_signed": str or null
}

If the value is illegible, use "[ILLEGIBLE]".
If the field is blank, value should be null and filled should be false.
'''

print('Form field extraction prompt defined.')
print('Handles: printed forms, questionnaires, applications.')

Dokumentklassifizierung vor der Extraktion

Schalten Sie vor der Extraktion einen Klassifizierungsschritt vor, um für jeden Dokumenttyp das passende Extraktionsschema anzuwenden:

import json

DOC_SCHEMAS = {
    'receipt': receipt_prompt,
    'form': form_prompt,
    'table': table_json_prompt,
    'letter': 'Extract all text preserving paragraph structure. Identify: sender, recipient, date, subject, body.',
    'label': 'Extract all text from this label. Include: product name, ingredients/contents, weight, expiry date, barcode numbers.'
}

def classify_and_extract(image_path):
    # Step 1: Classify document type
    classify_prompt = 'What type of document is this? Return JSON: {"type": "receipt|form|table|letter|label|other", "confidence": "high|medium|low"}'
    classification_text = extract_text(image_path, classify_prompt)
    doc_type = json.loads(classification_text)['type']

    # Step 2: Apply correct schema
    schema = DOC_SCHEMAS.get(doc_type, 'Extract all visible text from this document.')
    extracted = extract_text(image_path, schema)

    return {'type': doc_type, 'data': extracted}

print('Document classify-then-extract pipeline defined.')

Umgang mit Bildern geringer Qualität

Nicht alle Dokumentbilder sind klar. Prompts sollten auch mit Bildern schlechter Qualität zuverlässig umgehen:

low_quality_prompt = '''
Extract text from this document image. The image may be low quality, blurry, or poorly lit.

Extraction guidelines:
- Extract all text you can read with reasonable confidence
- For unclear sections, use [UNCLEAR] as a placeholder
- For completely unreadable sections, use [UNREADABLE: approximately N words]
- Do not guess or hallucinate words you cannot see clearly
- Note image quality issues at the end: "Image quality: [good/fair/poor]. Issues: [description]"

Be conservative — it is better to mark something as unclear than to guess incorrectly.
'''

print(low_quality_prompt)
print('\nConservative approach: unclear beats hallucinated.')

Zusammenfassung mehrseitiger Dokumente

Bei mehrseitigen Dokumenten, die als mehrere Bilder gesendet werden, kombinieren Sie die seitenweise Extraktion mit einem Syntheseschritt:

def extract_multi_page_document(image_paths):
    # Step 1: Extract text from each page
    page_texts = []
    for i, path in enumerate(image_paths):
        page_text = extract_text(path, f'Extract all text from page {i+1} of this document. Preserve structure.')
        page_texts.append(f'=== PAGE {i+1} ===\n{page_text}')

    full_text = '\n\n'.join(page_texts)

    # Step 2: Synthesize summary and key information
    r = client.messages.create(
        model='claude-opus-4-5', max_tokens=500,
        messages=[{'role': 'user', 'content': f'''
Here is the extracted text from a {len(image_paths)}-page document:\n\n{full_text}\n\n
Provide:
1. Document type and title
2. 3-sentence summary
3. Key data points extracted
Return JSON: {{"type": str, "title": str, "summary": str, "key_data": [str]}}
'''}]
    )
    return json.loads(r.content[0].text)

print('Multi-page document pipeline defined.')

Validierung nach der OCR-Extraktion

OCR-Ausgaben müssen vor ihrer Verwendung in nachgelagerten Systemen validiert werden. Häufige Validierungsprüfungen:

import re
from datetime import datetime

def validate_receipt_extraction(data):
    errors = []

    # Validate total is present and numeric
    if data.get('total') is None:
        errors.append('total is missing')
    elif not isinstance(data['total'], (int, float)):
        errors.append(f'total is not numeric: {data["total"]}')

    # Validate date format
    if data.get('transaction', {}).get('date'):
        date_str = data['transaction']['date']
        try:
            datetime.strptime(date_str, '%Y-%m-%d')
        except ValueError:
            errors.append(f'date format invalid: {date_str}')

    # Validate line items total approximately equals subtotal
    if data.get('items') and data.get('subtotal'):
        items_total = sum(item.get('total', 0) for item in data['items'] if item.get('total'))
        if abs(items_total - data['subtotal']) > 0.05:
            errors.append(f'Items total {items_total} does not match subtotal {data["subtotal"]}')

    return errors

print('Receipt validation function defined.')

Extraktion strukturierter Daten aus Diagrammen und Grafiken

Diagramme und Grafiken in Dokumentbildern enthalten Daten, die für herkömmliche OCR unsichtbar, für Vision-LLMs aber lesbar sind. Ein Prompt zur Diagrammextraktion fordert das Modell auf, die zugrunde liegenden Datenwerte abzulesen:

chart_prompt = '''
Extract the data from this chart or graph image.

Identify:
1. Chart type (bar, line, pie, scatter, table)
2. Title and axis labels
3. All data series names
4. All data points with their labels/values
5. Any notable trend or pattern

Return JSON:
{
  "chart_type": str,
  "title": str or null,
  "x_axis_label": str or null,
  "y_axis_label": str or null,
  "data_series": [
    {"name": str, "values": [{"label": str, "value": number}]}
  ],
  "key_insight": str
}

If exact values are not readable, provide best estimates with a note.
'''

print(chart_prompt)

Kurzer Check

Welches Vorgehen wird beim Transkribieren handschriftlicher Inhalte aus einem Bild für unleserliche Wörter empfohlen?

OCR und Dokumentanalyse – wichtigste Erkenntnisse

Vision-LLMs ermöglichen eine flexible Dokumentanalyse, die über die Zeichenerkennung hinausgeht:

  • Grundlegende Extraktion: Erhalten Sie Zeilenumbrüche und Absatzstruktur; geben Sie dies ausdrücklich an
  • Tabellen: Verwenden Sie das Markdown-Tabellenformat oder ein JSON-Schema mit Zeilen und Kopfzeilen, um die Struktur zu erhalten
  • Belege: Verwenden Sie ein eigenes Schema mit den Feldern für Händler, Positionen und Summen
  • Handschrift: Weisen Sie für Unleserliches [ILLEGIBLE] und für teilweise Lesbares [PARTIAL] an – raten Sie niemals
  • Formulare: Ordnen Sie Beschriftungen und Werte einander zu; geben Sie für jedes Feld den Status ausgefüllt/nicht ausgefüllt an
  • Klassifizieren Sie zuerst den Dokumenttyp und wenden Sie anschließend das passende Extraktionsschema an
  • Validieren Sie extrahierte Daten programmgesteuert: Pflichtfelder, numerische Typen, Datumsformate und mathematische Prüfungen
  • Bilder geringer Qualität: Eine konservative Extraktion ist besser als eine selbstsichere Halluzination
Kostenlos starten

Lerne AI Prompt Engineering mit einem KI-Tutor — kostenlos

Schreibe und führe echten Code in deinem Browser aus, bekomme sofortige Hilfe von einem 24/7 KI-Tutor und setze dein Lernen im Web oder in der App fort.

Kurse
53
Lektionen
199

Häufig gestellte Fragen

Ist die Lektion „Prompts für OCR und Dokumentenanalyse“ kostenlos?

Ja — der vollständige Text von „Prompts für OCR und Dokumentenanalyse“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des AI Prompt Engineering-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der AI Prompt Engineering-Kurs umfasst insgesamt 4 Lektionen.

Was lerne ich in „Prompts für OCR und Dokumentenanalyse“?

Extrahieren Sie Text, Tabellen und Strukturen aus Dokumentbildern Du übst AI Prompt Engineering mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.

Brauche ich Erfahrung, um AI Prompt Engineering zu starten?

Keine Vorkenntnisse erforderlich. AI Prompt Engineering auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 4 von 4.

Wie lange dauert die Lektion „Prompts für OCR und Dokumentenanalyse“?

Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.

Kann ich in dieser AI Prompt Engineering-Lektion Code schreiben und ausführen?

Ja. Jede AI Prompt Engineering-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.

Alle Lektionen in diesem Kurs

  1. Prompts für Bildbeschreibung und Captioning
  2. Visuelle Fragebeantwortung
  3. Prompts zum Vergleich mehrerer Bilder
  4. Prompts für OCR und Dokumentenanalyse
← Zurück zu AI Prompt Engineering