مطالبات OCR وتحليل المستندات
استخرجوا النصوص والجداول والبنية من صور المستندات
مطالبات OCR وتحليل المستندات درس مجاني في AI Prompt Engineering على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في AI Prompt Engineering، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة AI Prompt Engineering 4 دروس في المجموع.
نماذج LLMs بوصفها قارئات للمستندات
كان OCR (التعرّف البصري على الأحرف) يتطلب تقليديًا برامج متخصصة لاستخراج النص من الصور. أما نماذج LLMs للرؤية فأصبح بإمكانها الآن قراءة النص من الصور وفهم محتواه أيضًا — فهي لا تستخرج الأحرف فحسب، بل تحلل البنية والجداول والكتابة اليدوية والسياق.
مهام تحليل المستندات الشائعة:
- استخراج النص من المستندات الممسوحة ضوئيًا
- قراءة الإيصالات والفواتير والنماذج
- تحليل الجداول والمخططات
- نسخ الملاحظات المكتوبة بخط اليد
- قراءة الملصقات واللافتات المطبوعة
مطالبة استخراج النص الأساسي
لاستخراج نص بسيط من صورة مستند:
import anthropic, base64
client = anthropic.Anthropic(api_key='YOUR_API_KEY')
def extract_text(image_path, extraction_prompt):
with open(image_path, 'rb') as f:
img_b64 = base64.standard_b64encode(f.read()).decode('utf-8')
r = client.messages.create(
model='claude-opus-4-5', max_tokens=1000,
messages=[{'role': 'user', 'content': [
{'type': 'image', 'source': {'type': 'base64', 'media_type': 'image/jpeg', 'data': img_b64}},
{'type': 'text', 'text': extraction_prompt}
]}]
)
return r.content[0].text
# Basic extraction
basic_prompt = 'Extract all text from this document image exactly as it appears. Preserve line breaks.'
# Structure-preserving extraction
structured_prompt = 'Extract all text from this document image. Preserve: paragraph structure, line breaks, and any visible formatting. Do not add any text not present in the image.'
print('Text extraction functions defined.')الحفاظ على بنية الجدول
عندما يحتوي المستند على جدول، يؤدي استخراجه كنص مسطّح إلى فقدان بنيته. استخدم مطالبات تحافظ صراحةً على تنسيق الجدول:
table_prompt = '''
Extract all text from this document image.
If the document contains any tables, preserve the table structure using markdown table format:
| Column 1 | Column 2 | Column 3 |
|----------|----------|----------|
| Value | Value | Value |
For any text outside tables, use plain text preserving paragraph structure.
Do not invent or infer any data not visible in the image.
'''
# For structured output:
table_json_prompt = '''
Extract the table from this image.
Return JSON:
{
"headers": ["column name"],
"rows": [["cell value", "cell value"]],
"caption": "table caption if present or null"
}
If a cell is empty or illegible, use null.
'''
print('Table extraction prompts defined.')تفصيل عناصر الإيصال
للإيصالات بنية محددة — رأس (التاجر) وعناصر مفردة وإجماليات. وتستخرج مطالبة مخصصة للإيصالات هذه البنية بشكل موثوق:
import json
receipt_prompt = '''
Extract all information from this receipt image.
Return JSON:
{
"merchant": {
"name": str,
"address": str or null,
"phone": str or null
},
"transaction": {
"date": "YYYY-MM-DD or as written",
"time": "HH:MM or as written or null",
"receipt_number": str or null,
"payment_method": str or null
},
"items": [
{"description": str, "quantity": number or null, "unit_price": number or null, "total": number}
],
"subtotal": number or null,
"tax": number or null,
"tip": number or null,
"total": number,
"currency": "3-letter ISO code"
}
For any field not visible, use null. For numbers, use numeric type (not string).
'''
def extract_receipt(image_path):
text = extract_text(image_path, receipt_prompt)
return json.loads(text)
print('Receipt extraction function defined.')نسخ الملاحظات المكتوبة بخط اليد
يتطلب نسخ المحتوى المكتوب بخط اليد مطالبات تراعي التحديات، مثل الكلمات غير المقروءة والنص المشطوب والاختصارات:
handwriting_prompt = '''
Transcribe the handwritten text in this image as accurately as possible.
Handling rules:
- If a word is illegible, write [ILLEGIBLE]
- If a word is partially legible, write [PARTIAL: best_guess]
- If text is crossed out, include it with strikethrough notation: ~~crossed out text~~
- Preserve line breaks as they appear
- If there are arrows, circles, or annotations, note them in brackets: [arrow pointing right]
- Do not correct spelling or grammar
After transcription, estimate overall legibility: high (>90% readable) | medium (70-90%) | low (<70%)
Format:
TRANSCRIPTION:
[transcribed text here]
LEGIBILITY: [rating]
'''
print(handwriting_prompt)استخراج حقول النموذج
تحتوي النماذج المطبوعة على حقول معنونة وقيم مُدخلة. وتربط مطالبة استخراج النموذج بين التسميات والقيم:
form_prompt = '''
Extract all form fields and their values from this document image.
For each field:
- Field label: the printed label (e.g., "First Name:", "Date of Birth:")
- Field value: the filled-in value (handwritten or typed)
- Filled: whether the field has been filled in (true/false)
Return JSON:
{
"form_title": str or null,
"fields": [
{
"label": str,
"value": str or null,
"filled": true | false
}
],
"signature_present": true | false,
"date_signed": str or null
}
If the value is illegible, use "[ILLEGIBLE]".
If the field is blank, value should be null and filled should be false.
'''
print('Form field extraction prompt defined.')
print('Handles: printed forms, questionnaires, applications.')تصنيف المستند قبل الاستخراج
اربط خطوة تصنيف قبل الاستخراج لتطبيق مخطط الاستخراج الصحيح على كل نوع من المستندات:
import json
DOC_SCHEMAS = {
'receipt': receipt_prompt,
'form': form_prompt,
'table': table_json_prompt,
'letter': 'Extract all text preserving paragraph structure. Identify: sender, recipient, date, subject, body.',
'label': 'Extract all text from this label. Include: product name, ingredients/contents, weight, expiry date, barcode numbers.'
}
def classify_and_extract(image_path):
# Step 1: Classify document type
classify_prompt = 'What type of document is this? Return JSON: {"type": "receipt|form|table|letter|label|other", "confidence": "high|medium|low"}'
classification_text = extract_text(image_path, classify_prompt)
doc_type = json.loads(classification_text)['type']
# Step 2: Apply correct schema
schema = DOC_SCHEMAS.get(doc_type, 'Extract all visible text from this document.')
extracted = extract_text(image_path, schema)
return {'type': doc_type, 'data': extracted}
print('Document classify-then-extract pipeline defined.')التعامل مع الصور منخفضة الجودة
ليست كل صور المستندات واضحة. يجب أن تتعامل المطالبات مع الصور المتدهورة بطريقة متزنة:
low_quality_prompt = '''
Extract text from this document image. The image may be low quality, blurry, or poorly lit.
Extraction guidelines:
- Extract all text you can read with reasonable confidence
- For unclear sections, use [UNCLEAR] as a placeholder
- For completely unreadable sections, use [UNREADABLE: approximately N words]
- Do not guess or hallucinate words you cannot see clearly
- Note image quality issues at the end: "Image quality: [good/fair/poor]. Issues: [description]"
Be conservative — it is better to mark something as unclear than to guess incorrectly.
'''
print(low_quality_prompt)
print('\nConservative approach: unclear beats hallucinated.')تلخيص المستندات متعددة الصفحات
بالنسبة إلى المستندات متعددة الصفحات المرسلة كصور متعددة، اجمع بين الاستخراج صفحةً بصفحة وخطوة تركيبية:
def extract_multi_page_document(image_paths):
# Step 1: Extract text from each page
page_texts = []
for i, path in enumerate(image_paths):
page_text = extract_text(path, f'Extract all text from page {i+1} of this document. Preserve structure.')
page_texts.append(f'=== PAGE {i+1} ===\n{page_text}')
full_text = '\n\n'.join(page_texts)
# Step 2: Synthesize summary and key information
r = client.messages.create(
model='claude-opus-4-5', max_tokens=500,
messages=[{'role': 'user', 'content': f'''
Here is the extracted text from a {len(image_paths)}-page document:\n\n{full_text}\n\n
Provide:
1. Document type and title
2. 3-sentence summary
3. Key data points extracted
Return JSON: {{"type": str, "title": str, "summary": str, "key_data": [str]}}
'''}]
)
return json.loads(r.content[0].text)
print('Multi-page document pipeline defined.')التحقق بعد استخراج OCR
يجب التحقق من مخرجات OCR قبل استخدامها في الأنظمة اللاحقة. ومن فحوصات التحقق الشائعة:
import re
from datetime import datetime
def validate_receipt_extraction(data):
errors = []
# Validate total is present and numeric
if data.get('total') is None:
errors.append('total is missing')
elif not isinstance(data['total'], (int, float)):
errors.append(f'total is not numeric: {data["total"]}')
# Validate date format
if data.get('transaction', {}).get('date'):
date_str = data['transaction']['date']
try:
datetime.strptime(date_str, '%Y-%m-%d')
except ValueError:
errors.append(f'date format invalid: {date_str}')
# Validate line items total approximately equals subtotal
if data.get('items') and data.get('subtotal'):
items_total = sum(item.get('total', 0) for item in data['items'] if item.get('total'))
if abs(items_total - data['subtotal']) > 0.05:
errors.append(f'Items total {items_total} does not match subtotal {data["subtotal"]}')
return errors
print('Receipt validation function defined.')استخراج البيانات المنظّمة من المخططات والرسوم البيانية
تحتوي المخططات والرسوم البيانية في صور المستندات على بيانات لا يستطيع OCR التقليدي رؤيتها، لكن نماذج LLMs للرؤية تستطيع قراءتها. وتطلب مطالبة استخراج المخطط من النموذج قراءة قيم البيانات الأساسية:
chart_prompt = '''
Extract the data from this chart or graph image.
Identify:
1. Chart type (bar, line, pie, scatter, table)
2. Title and axis labels
3. All data series names
4. All data points with their labels/values
5. Any notable trend or pattern
Return JSON:
{
"chart_type": str,
"title": str or null,
"x_axis_label": str or null,
"y_axis_label": str or null,
"data_series": [
{"name": str, "values": [{"label": str, "value": number}]}
],
"key_insight": str
}
If exact values are not readable, provide best estimates with a note.
'''
print(chart_prompt)تحقق سريع
عند نسخ محتوى مكتوب بخط اليد من صورة، ما النهج الموصى به للكلمات غير المقروءة؟
OCR وتحليل المستندات — أهم النقاط
توفر نماذج LLMs للرؤية تحليلًا مرنًا للمستندات يتجاوز التعرّف على الأحرف:
- الاستخراج الأساسي: حافظ على فواصل الأسطر وبنية الفقرات، وحدد ذلك صراحةً
- الجداول: استخدم تنسيق جدول Markdown أو مخطط صفوف/رؤوس JSON للحفاظ على البنية
- الإيصالات: استخدم مخططًا مخصصًا يحتوي على حقول التاجر والعناصر والإجماليات
- الكتابة اليدوية: وجّه النموذج إلى استخدام [ILLEGIBLE] للنص غير المقروء و[PARTIAL] للنص المقروء جزئيًا — ولا تخمّن مطلقًا
- النماذج: اربط أزواج التسمية بالقيمة، وأدرج حالة التعبئة/عدم التعبئة لكل حقل
- صنّف نوع المستند أولًا، ثم طبّق مخطط الاستخراج المناسب
- تحقق من البيانات المستخرجة برمجيًا: الحقول المطلوبة، والأنواع الرقمية، وتنسيقات التاريخ، وفحوصات العمليات الحسابية
- الصور منخفضة الجودة: الاستخراج المتحفظ أفضل من الاختلاق الواثق
الأسئلة الشائعة
هل درس «مطالبات OCR وتحليل المستندات» مجاني؟
نعم — نص درس «مطالبات OCR وتحليل المستندات» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة AI Prompt Engineering، انتقل إلى CoddyKit PRO. تتضمن دورة AI Prompt Engineering 4 دروس في المجموع.
ماذا ستتعلم في «مطالبات OCR وتحليل المستندات»؟
استخرجوا النصوص والجداول والبنية من صور المستندات تتمرن على AI Prompt Engineering مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ AI Prompt Engineering؟
لا تُشترط خبرة سابقة. AI Prompt Engineering على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.
كم من الوقت يستغرق درس «مطالبات OCR وتحليل المستندات»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس AI Prompt Engineering هذا؟
نعم. كل درس في AI Prompt Engineering يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- مطالبات وصف الصور وإضافة التعليقات التوضيحية
- الإجابة عن الأسئلة المرئية
- مطالبات مقارنة صور متعددة
- مطالبات OCR وتحليل المستندات