الإجابة عن الأسئلة المرئية
اطلبوا إجابات عن أسئلة محددة حول محتوى الصورة وكمياته وسماته
الإجابة عن الأسئلة المرئية درس مجاني في AI Prompt Engineering على CoddyKit. هذا هو الدرس 2 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في AI Prompt Engineering، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة AI Prompt Engineering 4 دروس في المجموع.
الإجابة عن الأسئلة المرئية
الإجابة عن الأسئلة المرئية (VQA) هي مهمة الإجابة عن أسئلة باللغة الطبيعية حول صورة. وعلى خلاف وصف الصورة، الذي يصف كل شيء، يركز VQA النموذج على الإجابة عن سؤال محدد.
تكون مطالبات VQA دقيقة ومباشرة، وغالبًا ما تتطلب العد أو التعرّف أو المقارنة أو الاستدلال بشأن المحتوى المرئي. وتحدد جودة المطالبة ما إذا كنت ستحصل على إجابة دقيقة ومفيدة أم على رد عام ومبهم.
البنية الأساسية لمطالبة VQA
تجمع مطالبة VQA بين صورة وسؤال محدد. ويتمثل المفتاح في صياغة السؤال بدقة كافية لإنتاج إجابة مباشرة قابلة للاستخدام:
import anthropic, base64
client = anthropic.Anthropic(api_key='YOUR_API_KEY')
def ask_about_image(image_path, question, answer_format='Direct answer. No extra explanation.'):
with open(image_path, 'rb') as f:
img_b64 = base64.standard_b64encode(f.read()).decode('utf-8')
prompt = f'{question}\n\n{answer_format}'
r = client.messages.create(
model='claude-opus-4-5', max_tokens=150,
messages=[{'role': 'user', 'content': [
{'type': 'image', 'source': {'type': 'base64', 'media_type': 'image/jpeg', 'data': img_b64}},
{'type': 'text', 'text': prompt}
]}]
)
return r.content[0].text
# Example VQA calls (replace image.jpg with actual image)
print('VQA function defined. Ready for image questions.')أسئلة العدّ
يُعد العدّ من مهام VQA الشائعة. وتنتج مطالبات العدّ الدقيقة نتائج أكثر دقة من المطالبات المبهمة:
# Vague (bad):
vague_prompt = 'How many people are there?'
# Precise (good): specifies what counts and what does not
counting_prompt = '''
How many people are visible in this image?
Count only: people whose faces OR bodies are at least 50% visible.
Do NOT count: people who are heavily cropped, cut off at the edge, or only partially visible.
Return a single number.
'''
# Even more precise: handles partial visibility explicitly
precise_count = '''
Count the number of distinct individuals visible in this image.
If a person is partially obscured, count them if more than half their body is visible.
Return JSON: {"count": integer, "partially_visible": integer, "notes": "string or null"}
'''
print('Counting prompts: vague vs precise.')
print('Precise prompts define edge cases explicitly.')التعرّف على العلامات التجارية والشعارات
يُعد التعرّف على شعارات العلامات التجارية في الصور مهمة شائعة لتحليل المنتجات. ويجب أن تحدد المطالبة ما الذي ينبغي البحث عنه وتنسيق الإرجاع:
logo_prompt = '''
Identify all visible brand logos, company names, and product labels in this image.
For each, note:
- Brand/company name
- Where it appears in the image (top-left, center, on a product, etc.)
- Confidence: high (clearly legible) | medium (partially visible) | low (partially obscured)
Return JSON: {"brands": [{"name": str, "location": str, "confidence": str}]}
If no logos are visible, return: {"brands": []}
'''
import anthropic, base64, json
client = anthropic.Anthropic(api_key='YOUR_API_KEY')
def identify_brands(image_path):
with open(image_path, 'rb') as f:
img_b64 = base64.standard_b64encode(f.read()).decode('utf-8')
r = client.messages.create(
model='claude-opus-4-5', max_tokens=200,
messages=[{'role': 'user', 'content': [
{'type': 'image', 'source': {'type': 'base64', 'media_type': 'image/jpeg', 'data': img_b64}},
{'type': 'text', 'text': logo_prompt}
]}]
)
return json.loads(r.content[0].text)
print('Brand identification function defined.')التعرّف على المشاعر والتعبيرات
يتطلب التعرّف على التعبيرات العاطفية في الصور تصميم مطالبة دقيقًا يراعي عدم اليقين:
emotion_prompt = '''
Describe the emotional expression of the person in this image.
Assess:
- Primary emotion: (happy, sad, angry, surprised, fearful, disgusted, neutral, or other)
- Intensity: (low, moderate, high)
- Confidence: (high if expression is clear, medium if subtle, low if face is obscured or turned away)
- Evidence: which specific facial features support your assessment
Return JSON:
{
"primary_emotion": str,
"intensity": str,
"confidence": str,
"evidence": str,
"secondary_emotion": str or null
}
If no person or face is clearly visible, return: {"primary_emotion": null, "confidence": "none", "reason": str}
'''
print(emotion_prompt)أسئلة العلاقات المكانية
تتطلب الأسئلة حول مواضع العناصر بالنسبة إلى بعضها استخدام مفردات مكانية صريحة في المطالبة:
spatial_prompt = '''
Answer questions about the spatial relationships of objects in this image.
Use these spatial terms consistently:
- Position in frame: top-left, top-center, top-right, middle-left, center, middle-right, bottom-left, bottom-center, bottom-right
- Relative position: in front of, behind, to the left of, to the right of, above, below, overlapping
- Distance: in the foreground, in the midground, in the background
Question: {question}
Answer in one or two sentences using the spatial vocabulary above.
'''
# Example questions:
questions = [
'Where is the red cup relative to the laptop?',
'Is the plant in the foreground or background?',
'What object is to the left of the person?'
]
for q in questions:
print(spatial_prompt.replace('{question}', q)[:200])
print('---')تقييم الجودة والحالة
تقييم جودة العناصر أو حالتها في الصور — وهو مفيد لفحص المنتجات، وتقييم العقارات، ومراقبة الجودة:
condition_prompt = '''
Assess the condition of the main subject in this image.
Rate on these dimensions (1-5 scale, 5=excellent):
- Physical condition: (1=heavily damaged, 5=like new)
- Cleanliness: (1=very dirty, 5=spotless)
- Completeness: (1=major parts missing, 5=fully intact)
For each rating, provide one-sentence evidence.
Return JSON:
{
"physical_condition": {"score": int, "evidence": str},
"cleanliness": {"score": int, "evidence": str},
"completeness": {"score": int, "evidence": str},
"overall_grade": "excellent|good|fair|poor",
"recommendation": str
}
'''
print('Condition assessment prompt defined.')
print('Useful for: product inspection, real estate, equipment maintenance.')أسئلة VQA بنعم أو لا
تحتاج أسئلة نعم أو لا الثنائية إلى مطالبات تمنع النموذج من تقديم إجابة نثرية متحفظة عندما تحتاج إلى قيمة منطقية بسيطة:
def yes_no_question(image_path, question):
with open(image_path, 'rb') as f:
img_b64 = base64.standard_b64encode(f.read()).decode('utf-8')
prompt = f'''
Answer this yes/no question about the image.
Return JSON: {{"answer": "yes|no", "confidence": "high|medium|low", "reason": str}}
Do NOT answer with maybe, possibly, or a hedged statement.
If you genuinely cannot determine the answer, return {{"answer": "unclear", "confidence": "low", "reason": str}}
Question: {question}
'''
r = client.messages.create(
model='claude-opus-4-5', max_tokens=100,
messages=[{'role': 'user', 'content': [
{'type': 'image', 'source': {'type': 'base64', 'media_type': 'image/jpeg', 'data': img_b64}},
{'type': 'text', 'text': prompt}
]}]
)
return json.loads(r.content[0].text)
# Example: 'Is there a safety helmet visible in the image?'
print('Yes/no VQA function defined.')تسلسل أسئلة VQA
يمكن تسلسل عدة أسئلة VQA حول الصورة نفسها في مطالبة واحدة لتقليل عدد استدعاءات API:
multi_question_prompt = '''
Answer all of the following questions about this image.
Return a JSON object where each key is the question ID.
Questions:
1. How many people are visible?
2. What is the approximate age range of the youngest person?
3. Is there any food visible in the image?
4. What is the dominant color in the image?
5. Is the setting indoors or outdoors?
Return JSON:
{
"q1": {"answer": str},
"q2": {"answer": str},
"q3": {"answer": "yes|no", "details": str or null},
"q4": {"answer": str},
"q5": {"answer": "indoors|outdoors|unclear"}
}
'''
print('Multi-question VQA prompt — answers 5 questions in one API call.')التعامل مع عدم اليقين في VQA
لا يمكن أحيانًا الإجابة عن أسئلة VQA بيقين — فقد تكون الصورة ضبابية، أو يكون العنصر المعني محجوبًا جزئيًا، أو تكون الإجابة ملتبسة فعلًا. اطلب إظهار عدم اليقين صراحةً بدلًا من إجبار النموذج على التخمين:
uncertainty_vqa_prompt = '''
Answer this question about the image as precisely as possible.
If the answer is not clearly visible or is ambiguous, say so explicitly.
Question: {question}
Return JSON:
{
"answer": str,
"confidence": "high|medium|low|cannot_determine",
"limitation": str or null
}
For confidence levels:
- high: Answer is clearly visible and unambiguous
- medium: Visible but some uncertainty
- low: Partially visible or requires inference
- cannot_determine: Not enough visual information
Question: What brand is printed on the water bottle?
'''
print(uncertainty_vqa_prompt)مطالبات VQA الخاصة بالمجال
تتطلب المجالات المختلفة مفردات VQA ومعايير قياس مختلفة. وتنتج المطالبات الخاصة بالمجال إجابات أدق وأكثر قابلية للتنفيذ:
# Manufacturing quality control VQA
qc_prompt = '''
Inspect this product image for quality defects.
Answer each question:
1. Are there any visible scratches or surface damage? (yes/no + location)
2. Is the product alignment within expected tolerance? (yes/no)
3. Are all required labels/markings present? (yes/no + list missing ones)
4. Overall QC result: PASS or FAIL?
Return JSON:
{"scratches": {"present": bool, "location": str or null},
"alignment_ok": bool,
"labels_complete": bool, "missing_labels": [str],
"qc_result": "PASS|FAIL",
"fail_reasons": [str]}
'''
# Food safety VQA
food_prompt = '''
Inspect this food preparation image.
1. Are gloves being worn? 2. Is hair covered? 3. Any visible contamination risk?
Return JSON: {"gloves": bool, "hair_covered": bool, "contamination_risk": bool, "details": str}
'''
print("Domain-specific QC and food safety VQA prompts defined.")تحقق سريع
أي مطالبة VQA يُرجح أن تنتج إجابة دقيقة وقابلة للاستخدام عند عدّ العناصر في صورة؟
مطالبات VQA — أهم النقاط
تتطلب الإجابة الفعّالة عن الأسئلة المرئية تصميم مطالبات دقيقة:
- اجعل الأسئلة محددة ومباشرة — وتجنب المصطلحات المبهمة مثل بعض أو متنوعة
- حدد الحالات الحدّية صراحةً في أسئلة العدّ (ما الذي يُعدّ ظاهرًا جزئيًا؟)
- حدد تنسيق الإخراج بدقة — JSON أو رقم واحد أو نعم/لا — لمنع الإجابات النثرية المترددة
- أدرج مستويات الثقة لجميع الإجابات حتى يمكن تمييز النتائج غير المؤكدة
- اجمع عدة أسئلة حول الصورة نفسها في استدعاء واحد لتقليل تكاليف API
- في أسئلة نعم/لا الثنائية، امنع الإجابات المترددة صراحةً، مع توفير خيار غير واضح
- تتطلب VQA المتخصصة بمجال معين (الطب أو القانون أو المنتجات) استخدام مفردات المجال في المطالبة
الأسئلة الشائعة
هل درس «الإجابة عن الأسئلة المرئية» مجاني؟
نعم — نص درس «الإجابة عن الأسئلة المرئية» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة AI Prompt Engineering، انتقل إلى CoddyKit PRO. تتضمن دورة AI Prompt Engineering 4 دروس في المجموع.
ماذا ستتعلم في «الإجابة عن الأسئلة المرئية»؟
اطلبوا إجابات عن أسئلة محددة حول محتوى الصورة وكمياته وسماته تتمرن على AI Prompt Engineering مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ AI Prompt Engineering؟
لا تُشترط خبرة سابقة. AI Prompt Engineering على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 2 من أصل 4.
كم من الوقت يستغرق درس «الإجابة عن الأسئلة المرئية»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس AI Prompt Engineering هذا؟
نعم. كل درس في AI Prompt Engineering يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- مطالبات وصف الصور وإضافة التعليقات التوضيحية
- الإجابة عن الأسئلة المرئية
- مطالبات مقارنة صور متعددة
- مطالبات OCR وتحليل المستندات