0Pricing
AI Prompt Engineering · Ders

Görsel Soru Yanıtlama

Görüntü içeriği, miktarlar ve nitelikler hakkında özgül sorular sorun.

Görsel Soru Yanıtlama, CoddyKit'te ücretsiz bir AI Prompt Engineering dersidir. Bu, 4 dersinin 2. dersidir. Aşağıdan dersin tamamını ücretsiz okuyabilir, sonra tarayıcıda yerleşik kod editörü ve 7/24 yapay zeka koçu ile uygulamalı olarak pratik yapabilirsin. Bu, AI Prompt Engineering öğrenme yolunun bir parçasıdır ve ilerlemeniz web ve CoddyKit uygulaması arasında senkronize olur. AI Prompt Engineering kursu toplamda 4 dersten oluşur.

Görsel Soru Yanıtlama

Görsel Soru Yanıtlama (VQA), bir görüntü hakkındaki doğal dil sorularını yanıtlama görevidir. Her şeyi açıklayan görüntü açıklamasından farklı olarak VQA, modeli belirli bir soruyu yanıtlamaya odaklar.

VQA istemleri kesin ve doğrudandır; genellikle görsel içeriği saymayı, tanımlamayı, karşılaştırmayı veya onun hakkında akıl yürütmeyi gerektirir. İstemin kalitesi, kesin ve kullanışlı bir yanıt mı yoksa muğlak, genel bir yanıt mı alacağınızı belirler.

Temel VQA İstemi Yapısı

Bir VQA istemi, görüntüyü belirli bir soruyla eşleştirir. Önemli olan, doğrudan ve kullanılabilir bir yanıt üretecek kadar kesin bir soru sormaktır:

import anthropic, base64

client = anthropic.Anthropic(api_key='YOUR_API_KEY')

def ask_about_image(image_path, question, answer_format='Direct answer. No extra explanation.'):
    with open(image_path, 'rb') as f:
        img_b64 = base64.standard_b64encode(f.read()).decode('utf-8')

    prompt = f'{question}\n\n{answer_format}'

    r = client.messages.create(
        model='claude-opus-4-5', max_tokens=150,
        messages=[{'role': 'user', 'content': [
            {'type': 'image', 'source': {'type': 'base64', 'media_type': 'image/jpeg', 'data': img_b64}},
            {'type': 'text', 'text': prompt}
        ]}]
    )
    return r.content[0].text

# Example VQA calls (replace image.jpg with actual image)
print('VQA function defined. Ready for image questions.')

Sayma Soruları

Sayma, yaygın bir VQA görevidir. Kesin sayma istemleri, muğlak istemlere kıyasla daha doğru sonuçlar üretir:

# Vague (bad):
vague_prompt = 'How many people are there?'

# Precise (good): specifies what counts and what does not
counting_prompt = '''
How many people are visible in this image?
Count only: people whose faces OR bodies are at least 50% visible.
Do NOT count: people who are heavily cropped, cut off at the edge, or only partially visible.
Return a single number.
'''

# Even more precise: handles partial visibility explicitly
precise_count = '''
Count the number of distinct individuals visible in this image.
If a person is partially obscured, count them if more than half their body is visible.
Return JSON: {"count": integer, "partially_visible": integer, "notes": "string or null"}
'''

print('Counting prompts: vague vs precise.')
print('Precise prompts define edge cases explicitly.')

Marka ve Logo Tanıma

Görüntülerde marka logolarını tanımak, yaygın bir ürün analizi görevidir. İstem, ne aranacağını ve hangi biçimde döndürüleceğini belirtmelidir:

logo_prompt = '''
Identify all visible brand logos, company names, and product labels in this image.

For each, note:
- Brand/company name
- Where it appears in the image (top-left, center, on a product, etc.)
- Confidence: high (clearly legible) | medium (partially visible) | low (partially obscured)

Return JSON: {"brands": [{"name": str, "location": str, "confidence": str}]}
If no logos are visible, return: {"brands": []}
'''

import anthropic, base64, json
client = anthropic.Anthropic(api_key='YOUR_API_KEY')

def identify_brands(image_path):
    with open(image_path, 'rb') as f:
        img_b64 = base64.standard_b64encode(f.read()).decode('utf-8')
    r = client.messages.create(
        model='claude-opus-4-5', max_tokens=200,
        messages=[{'role': 'user', 'content': [
            {'type': 'image', 'source': {'type': 'base64', 'media_type': 'image/jpeg', 'data': img_b64}},
            {'type': 'text', 'text': logo_prompt}
        ]}]
    )
    return json.loads(r.content[0].text)

print('Brand identification function defined.')

Duygu ve İfade Tanıma

Görüntülerdeki duygusal ifadeleri tanımak, belirsizliği kabul eden dikkatli bir istem tasarımı gerektirir:

emotion_prompt = '''
Describe the emotional expression of the person in this image.

Assess:
- Primary emotion: (happy, sad, angry, surprised, fearful, disgusted, neutral, or other)
- Intensity: (low, moderate, high)
- Confidence: (high if expression is clear, medium if subtle, low if face is obscured or turned away)
- Evidence: which specific facial features support your assessment

Return JSON:
{
  "primary_emotion": str,
  "intensity": str,
  "confidence": str,
  "evidence": str,
  "secondary_emotion": str or null
}

If no person or face is clearly visible, return: {"primary_emotion": null, "confidence": "none", "reason": str}
'''

print(emotion_prompt)

Mekânsal İlişki Soruları

Nesnelerin birbirine göre nerede olduğuna ilişkin sorular, istemde açık mekânsal sözcük dağarcığı gerektirir:

spatial_prompt = '''
Answer questions about the spatial relationships of objects in this image.
Use these spatial terms consistently:
- Position in frame: top-left, top-center, top-right, middle-left, center, middle-right, bottom-left, bottom-center, bottom-right
- Relative position: in front of, behind, to the left of, to the right of, above, below, overlapping
- Distance: in the foreground, in the midground, in the background

Question: {question}

Answer in one or two sentences using the spatial vocabulary above.
'''

# Example questions:
questions = [
    'Where is the red cup relative to the laptop?',
    'Is the plant in the foreground or background?',
    'What object is to the left of the person?'
]

for q in questions:
    print(spatial_prompt.replace('{question}', q)[:200])
    print('---')

Kalite ve Durum Değerlendirmesi

Görüntülerdeki nesnelerin kalitesini veya durumunu değerlendirme — ürün incelemesi, gayrimenkul değerlendirmesi ve kalite kontrolü için kullanışlıdır:

condition_prompt = '''
Assess the condition of the main subject in this image.

Rate on these dimensions (1-5 scale, 5=excellent):
- Physical condition: (1=heavily damaged, 5=like new)
- Cleanliness: (1=very dirty, 5=spotless)
- Completeness: (1=major parts missing, 5=fully intact)

For each rating, provide one-sentence evidence.

Return JSON:
{
  "physical_condition": {"score": int, "evidence": str},
  "cleanliness": {"score": int, "evidence": str},
  "completeness": {"score": int, "evidence": str},
  "overall_grade": "excellent|good|fair|poor",
  "recommendation": str
}
'''

print('Condition assessment prompt defined.')
print('Useful for: product inspection, real estate, equipment maintenance.')

Evet/Hayır VQA Soruları

İkili evet/hayır soruları, basit bir mantıksal değere ihtiyaç duyduğunuzda modelin temkinli, düzyazı bir yanıt vermesini önleyen istemler gerektirir:

def yes_no_question(image_path, question):
    with open(image_path, 'rb') as f:
        img_b64 = base64.standard_b64encode(f.read()).decode('utf-8')

    prompt = f'''
Answer this yes/no question about the image.
Return JSON: {{"answer": "yes|no", "confidence": "high|medium|low", "reason": str}}
Do NOT answer with maybe, possibly, or a hedged statement.
If you genuinely cannot determine the answer, return {{"answer": "unclear", "confidence": "low", "reason": str}}

Question: {question}
'''

    r = client.messages.create(
        model='claude-opus-4-5', max_tokens=100,
        messages=[{'role': 'user', 'content': [
            {'type': 'image', 'source': {'type': 'base64', 'media_type': 'image/jpeg', 'data': img_b64}},
            {'type': 'text', 'text': prompt}
        ]}]
    )
    return json.loads(r.content[0].text)

# Example: 'Is there a safety helmet visible in the image?'
print('Yes/no VQA function defined.')

VQA Sorularını Zincirleme

Aynı görüntü hakkındaki birden çok VQA sorusu, API çağrılarını azaltmak için tek bir istemde zincirlenebilir:

multi_question_prompt = '''
Answer all of the following questions about this image.
Return a JSON object where each key is the question ID.

Questions:
1. How many people are visible?
2. What is the approximate age range of the youngest person?
3. Is there any food visible in the image?
4. What is the dominant color in the image?
5. Is the setting indoors or outdoors?

Return JSON:
{
  "q1": {"answer": str},
  "q2": {"answer": str},
  "q3": {"answer": "yes|no", "details": str or null},
  "q4": {"answer": str},
  "q5": {"answer": "indoors|outdoors|unclear"}
}
'''

print('Multi-question VQA prompt — answers 5 questions in one API call.')

VQA Belirsizliğini Ele Alma

VQA soruları bazen kesin olarak yanıtlanamaz — görüntü bulanık olabilir, ilgili unsur kısmen gizlenmiş olabilir veya yanıt gerçekten muğlak olabilir. Bir tahmini zorlamak yerine açıkça belirsizlik belirtmesini isteyin:

uncertainty_vqa_prompt = '''
Answer this question about the image as precisely as possible.
If the answer is not clearly visible or is ambiguous, say so explicitly.

Question: {question}

Return JSON:
{
  "answer": str,
  "confidence": "high|medium|low|cannot_determine",
  "limitation": str or null
}

For confidence levels:
- high: Answer is clearly visible and unambiguous
- medium: Visible but some uncertainty
- low: Partially visible or requires inference
- cannot_determine: Not enough visual information

Question: What brand is printed on the water bottle?
'''

print(uncertainty_vqa_prompt)

Alana Özgü VQA İstemleri

Farklı alanlar, farklı VQA kelime dağarcığı ve ölçüm standartları gerektirir. Alana özgü istemler daha doğru ve eyleme dönüştürülebilir yanıtlar üretir:

# Manufacturing quality control VQA
qc_prompt = '''
Inspect this product image for quality defects.
Answer each question:
1. Are there any visible scratches or surface damage? (yes/no + location)
2. Is the product alignment within expected tolerance? (yes/no)
3. Are all required labels/markings present? (yes/no + list missing ones)
4. Overall QC result: PASS or FAIL?

Return JSON:
{"scratches": {"present": bool, "location": str or null},
 "alignment_ok": bool,
 "labels_complete": bool, "missing_labels": [str],
 "qc_result": "PASS|FAIL",
 "fail_reasons": [str]}
'''

# Food safety VQA  
food_prompt = '''
Inspect this food preparation image.
1. Are gloves being worn? 2. Is hair covered? 3. Any visible contamination risk?
Return JSON: {"gloves": bool, "hair_covered": bool, "contamination_risk": bool, "details": str}
'''

print("Domain-specific QC and food safety VQA prompts defined.")

Hızlı Kontrol

Görüntüdeki nesneleri sayarken hangi VQA isteminin kesin ve kullanılabilir bir yanıt üretme olasılığı en yüksektir?

VQA İstemleri — Önemli Noktalar

Etkili Görsel Soru Yanıtlama, özenle tasarlanmış istemler gerektirir:

  • Soruları özgül ve doğrudan ifade edin — bazı veya çeşitli gibi belirsiz ifadelerden kaçının
  • Sayma sorularında sınır durumlarını açıkça tanımlayın (kısmen görünen bir şey ne zaman sayılır?)
  • Kaçamak düzyazı yanıtlarını önlemek için kesin çıktı biçimini belirtin — JSON, tek sayı, evet/hayır
  • Belirsiz yanıtların işaretlenebilmesi için tüm yanıtlar için güven düzeyleri ekleyin
  • API maliyetlerini azaltmak için aynı görüntü hakkındaki birden çok soruyu tek bir çağrıda gruplayın
  • İkili evet/hayır sorularında, belirsiz seçeneğini açıkça bir kaçış seçeneği olarak tanımlayarak kaçamak yanıtları engelleyin
  • Alana özgü VQA (tıp, hukuk, ürün) istemde alan terminolojisi gerektirir

Sıkça Sorulan Sorular

“Görsel Soru Yanıtlama” dersi ücretsiz mi?

Evet — “Görsel Soru Yanıtlama” dersin tüm metni burada web'de ücretsiz olarak okunabilir. Etkileşimli olarak pratik yapmak (yerleşik kod editörü ve 7/24 yapay zeka koçu) ve AI Prompt Engineering kursunun geri kalanını açmak için CoddyKit PRO'ya yükselt. AI Prompt Engineering kursu toplamda 4 dersten oluşur.

“Görsel Soru Yanıtlama” dersinde ne öğreneceğim?

Görüntü içeriği, miktarlar ve nitelikler hakkında özgül sorular sorun. AI Prompt Engineering ile uygulamalı kodu tarayıcıda doğrudan çalıştırarak pratik yaparsın ve 7/24 yapay zeka koçu dersi çalışırken sorularını yanıtlar.

AI Prompt Engineering öğrenmeye başlamak için deneyim gerekli mi?

Önceden deneyim gerekmez. CoddyKit'te AI Prompt Engineering, başlangıçtan ileri seviyeye kadar yapılandırıldığı için buradan başlayabilir veya başından başlayıp kendi hızında ilerleme yapabilirsin. Bu, 4 dersinin 2. dersidir.

“Görsel Soru Yanıtlama” dersi ne kadar sürer?

Çoğu CoddyKit dersi yaklaşık 5–10 dakika sürer. Her biri kısa ve etkileşimli olduğu için sabit ilerleme yaparsın ve web ile uygulama arasında tam olarak bıraktığın yerden devam edebilirsin.

Bu AI Prompt Engineering dersinde kod yazıp çalıştırabilir miyim?

Evet. Her AI Prompt Engineering dersi yerleşik bir kod editörü içerir, bu sayede tarayıcıda gerçek kod yazıp çalıştırabilir ve anlık yapay zeka geri bildirimi alırsın — yerel kurulum gerekli değildir.

Bu kursun tüm dersleri

  1. Görüntü Açıklama ve Altyazı İstemleri
  2. Görsel Soru Yanıtlama
  3. Birden Çok Görüntüyü Karşılaştırma İstemleri
  4. OCR ve Belge Analizi İstemleri
← AI Prompt Engineering Sayfasına Dön