0Pricing
AI Prompt Engineering · Pelajaran

Penjawaban Pertanyaan Visual

Ajukan pertanyaan spesifik tentang isi, jumlah, dan atribut gambar.

Penjawaban Pertanyaan Visual adalah pelajaran AI Prompt Engineering gratis di CoddyKit. Ini adalah pelajaran 2 dari 4. Kamu bisa membaca pelajaran lengkapnya di bawah secara gratis — lalu praktikkan langsung di browser dengan editor kode bawaan dan tutor AI 24/7. Ini adalah bagian dari jalur belajar AI Prompt Engineering, dan progresmu tersinkronisasi di web dan aplikasi CoddyKit. Kursus AI Prompt Engineering mencakup 4 pelajaran total.

Penjawaban Pertanyaan Visual

Penjawaban Pertanyaan Visual (VQA) adalah tugas menjawab pertanyaan dalam bahasa alami tentang sebuah gambar. Berbeda dari deskripsi gambar (yang mendeskripsikan semuanya), VQA memfokuskan model untuk menjawab pertanyaan tertentu.

Prompt VQA bersifat tepat dan langsung, serta sering memerlukan penghitungan, identifikasi, perbandingan, atau penalaran tentang konten visual. Kualitas prompt menentukan apakah Anda mendapatkan jawaban yang tepat dan bermanfaat atau respons umum yang samar.

Struktur Dasar Prompt VQA

Prompt VQA memasangkan gambar dengan pertanyaan tertentu. Kuncinya adalah membuat pertanyaan cukup tepat agar menghasilkan jawaban langsung yang dapat digunakan:

import anthropic, base64

client = anthropic.Anthropic(api_key='YOUR_API_KEY')

def ask_about_image(image_path, question, answer_format='Direct answer. No extra explanation.'):
    with open(image_path, 'rb') as f:
        img_b64 = base64.standard_b64encode(f.read()).decode('utf-8')

    prompt = f'{question}\n\n{answer_format}'

    r = client.messages.create(
        model='claude-opus-4-5', max_tokens=150,
        messages=[{'role': 'user', 'content': [
            {'type': 'image', 'source': {'type': 'base64', 'media_type': 'image/jpeg', 'data': img_b64}},
            {'type': 'text', 'text': prompt}
        ]}]
    )
    return r.content[0].text

# Example VQA calls (replace image.jpg with actual image)
print('VQA function defined. Ready for image questions.')

Pertanyaan Penghitungan

Menghitung adalah tugas VQA yang umum. Prompt penghitungan yang tepat menghasilkan hasil yang lebih akurat daripada prompt yang samar:

# Vague (bad):
vague_prompt = 'How many people are there?'

# Precise (good): specifies what counts and what does not
counting_prompt = '''
How many people are visible in this image?
Count only: people whose faces OR bodies are at least 50% visible.
Do NOT count: people who are heavily cropped, cut off at the edge, or only partially visible.
Return a single number.
'''

# Even more precise: handles partial visibility explicitly
precise_count = '''
Count the number of distinct individuals visible in this image.
If a person is partially obscured, count them if more than half their body is visible.
Return JSON: {"count": integer, "partially_visible": integer, "notes": "string or null"}
'''

print('Counting prompts: vague vs precise.')
print('Precise prompts define edge cases explicitly.')

Identifikasi Merek dan Logo

Mengidentifikasi logo merek dalam gambar adalah tugas umum dalam analisis produk. Prompt harus menentukan hal yang perlu dicari dan format keluarannya:

logo_prompt = '''
Identify all visible brand logos, company names, and product labels in this image.

For each, note:
- Brand/company name
- Where it appears in the image (top-left, center, on a product, etc.)
- Confidence: high (clearly legible) | medium (partially visible) | low (partially obscured)

Return JSON: {"brands": [{"name": str, "location": str, "confidence": str}]}
If no logos are visible, return: {"brands": []}
'''

import anthropic, base64, json
client = anthropic.Anthropic(api_key='YOUR_API_KEY')

def identify_brands(image_path):
    with open(image_path, 'rb') as f:
        img_b64 = base64.standard_b64encode(f.read()).decode('utf-8')
    r = client.messages.create(
        model='claude-opus-4-5', max_tokens=200,
        messages=[{'role': 'user', 'content': [
            {'type': 'image', 'source': {'type': 'base64', 'media_type': 'image/jpeg', 'data': img_b64}},
            {'type': 'text', 'text': logo_prompt}
        ]}]
    )
    return json.loads(r.content[0].text)

print('Brand identification function defined.')

Pengenalan Emosi dan Ekspresi

Mengidentifikasi ekspresi emosional dalam gambar memerlukan perancangan prompt yang cermat dan mengakui adanya ketidakpastian:

emotion_prompt = '''
Describe the emotional expression of the person in this image.

Assess:
- Primary emotion: (happy, sad, angry, surprised, fearful, disgusted, neutral, or other)
- Intensity: (low, moderate, high)
- Confidence: (high if expression is clear, medium if subtle, low if face is obscured or turned away)
- Evidence: which specific facial features support your assessment

Return JSON:
{
  "primary_emotion": str,
  "intensity": str,
  "confidence": str,
  "evidence": str,
  "secondary_emotion": str or null
}

If no person or face is clearly visible, return: {"primary_emotion": null, "confidence": "none", "reason": str}
'''

print(emotion_prompt)

Pertanyaan tentang Hubungan Spasial

Pertanyaan tentang posisi objek relatif terhadap satu sama lain memerlukan kosakata spasial yang eksplisit dalam prompt:

spatial_prompt = '''
Answer questions about the spatial relationships of objects in this image.
Use these spatial terms consistently:
- Position in frame: top-left, top-center, top-right, middle-left, center, middle-right, bottom-left, bottom-center, bottom-right
- Relative position: in front of, behind, to the left of, to the right of, above, below, overlapping
- Distance: in the foreground, in the midground, in the background

Question: {question}

Answer in one or two sentences using the spatial vocabulary above.
'''

# Example questions:
questions = [
    'Where is the red cup relative to the laptop?',
    'Is the plant in the foreground or background?',
    'What object is to the left of the person?'
]

for q in questions:
    print(spatial_prompt.replace('{question}', q)[:200])
    print('---')

Penilaian Kualitas dan Kondisi

Menilai kualitas atau kondisi objek dalam gambar—berguna untuk pemeriksaan produk, penilaian real estat, dan pengendalian kualitas:

condition_prompt = '''
Assess the condition of the main subject in this image.

Rate on these dimensions (1-5 scale, 5=excellent):
- Physical condition: (1=heavily damaged, 5=like new)
- Cleanliness: (1=very dirty, 5=spotless)
- Completeness: (1=major parts missing, 5=fully intact)

For each rating, provide one-sentence evidence.

Return JSON:
{
  "physical_condition": {"score": int, "evidence": str},
  "cleanliness": {"score": int, "evidence": str},
  "completeness": {"score": int, "evidence": str},
  "overall_grade": "excellent|good|fair|poor",
  "recommendation": str
}
'''

print('Condition assessment prompt defined.')
print('Useful for: product inspection, real estate, equipment maintenance.')

Pertanyaan VQA Ya/Tidak

Pertanyaan biner ya/tidak memerlukan prompt yang mencegah model memberikan jawaban prosa yang penuh keraguan saat Anda membutuhkan boolean sederhana:

def yes_no_question(image_path, question):
    with open(image_path, 'rb') as f:
        img_b64 = base64.standard_b64encode(f.read()).decode('utf-8')

    prompt = f'''
Answer this yes/no question about the image.
Return JSON: {{"answer": "yes|no", "confidence": "high|medium|low", "reason": str}}
Do NOT answer with maybe, possibly, or a hedged statement.
If you genuinely cannot determine the answer, return {{"answer": "unclear", "confidence": "low", "reason": str}}

Question: {question}
'''

    r = client.messages.create(
        model='claude-opus-4-5', max_tokens=100,
        messages=[{'role': 'user', 'content': [
            {'type': 'image', 'source': {'type': 'base64', 'media_type': 'image/jpeg', 'data': img_b64}},
            {'type': 'text', 'text': prompt}
        ]}]
    )
    return json.loads(r.content[0].text)

# Example: 'Is there a safety helmet visible in the image?'
print('Yes/no VQA function defined.')

Merangkai Pertanyaan VQA

Beberapa pertanyaan VQA tentang gambar yang sama dapat dirangkai dalam satu prompt untuk mengurangi panggilan API:

multi_question_prompt = '''
Answer all of the following questions about this image.
Return a JSON object where each key is the question ID.

Questions:
1. How many people are visible?
2. What is the approximate age range of the youngest person?
3. Is there any food visible in the image?
4. What is the dominant color in the image?
5. Is the setting indoors or outdoors?

Return JSON:
{
  "q1": {"answer": str},
  "q2": {"answer": str},
  "q3": {"answer": "yes|no", "details": str or null},
  "q4": {"answer": str},
  "q5": {"answer": "indoors|outdoors|unclear"}
}
'''

print('Multi-question VQA prompt — answers 5 questions in one API call.')

Menangani Ketidakpastian VQA

Pertanyaan VQA terkadang tidak dapat dijawab dengan pasti—gambar mungkin buram, elemen yang relevan mungkin sebagian tertutup, atau jawabannya memang ambigu. Minta ketidakpastian secara eksplisit alih-alih memaksa tebakan:

uncertainty_vqa_prompt = '''
Answer this question about the image as precisely as possible.
If the answer is not clearly visible or is ambiguous, say so explicitly.

Question: {question}

Return JSON:
{
  "answer": str,
  "confidence": "high|medium|low|cannot_determine",
  "limitation": str or null
}

For confidence levels:
- high: Answer is clearly visible and unambiguous
- medium: Visible but some uncertainty
- low: Partially visible or requires inference
- cannot_determine: Not enough visual information

Question: What brand is printed on the water bottle?
'''

print(uncertainty_vqa_prompt)

Prompt VQA Khusus Domain

Domain yang berbeda memerlukan kosakata dan standar pengukuran VQA yang berbeda. Prompt khusus domain menghasilkan jawaban yang lebih akurat dan dapat ditindaklanjuti:

# Manufacturing quality control VQA
qc_prompt = '''
Inspect this product image for quality defects.
Answer each question:
1. Are there any visible scratches or surface damage? (yes/no + location)
2. Is the product alignment within expected tolerance? (yes/no)
3. Are all required labels/markings present? (yes/no + list missing ones)
4. Overall QC result: PASS or FAIL?

Return JSON:
{"scratches": {"present": bool, "location": str or null},
 "alignment_ok": bool,
 "labels_complete": bool, "missing_labels": [str],
 "qc_result": "PASS|FAIL",
 "fail_reasons": [str]}
'''

# Food safety VQA  
food_prompt = '''
Inspect this food preparation image.
1. Are gloves being worn? 2. Is hair covered? 3. Any visible contamination risk?
Return JSON: {"gloves": bool, "hair_covered": bool, "contamination_risk": bool, "details": str}
'''

print("Domain-specific QC and food safety VQA prompts defined.")

Pemeriksaan Cepat

Prompt VQA manakah yang paling mungkin menghasilkan jawaban yang tepat dan dapat digunakan saat menghitung objek dalam gambar?

VQA — Poin-Poin Utama

Visual Question Answering yang efektif memerlukan perancangan prompt yang sangat tepat:

  • Buat pertanyaan yang spesifik dan langsung — hindari istilah samar seperti beberapa atau beragam
  • Tentukan kasus tepi secara eksplisit untuk pertanyaan penghitungan (apa yang dianggap terlihat sebagian?)
  • Tentukan format output yang tepat — JSON, satu angka, ya/tidak — untuk mencegah jawaban berupa uraian yang bersifat mendua
  • Sertakan tingkat keyakinan untuk semua jawaban agar output yang tidak pasti dapat ditandai
  • Gabungkan beberapa pertanyaan tentang gambar yang sama dalam satu panggilan untuk mengurangi biaya API
  • Untuk pertanyaan biner ya/tidak, cegah respons mendua secara eksplisit dengan menyediakan pilihan tidak jelas
  • VQA khusus domain (medis, hukum, produk) memerlukan kosakata domain di dalam prompt

Pertanyaan yang Sering Diajukan

Apakah pelajaran “Penjawaban Pertanyaan Visual” gratis?

Ya — teks lengkap “Penjawaban Pertanyaan Visual” gratis dibaca di sini di web. Untuk praktiknya secara interaktif (editor kode bawaan dan tutor AI 24/7) dan buka sisa kursus AI Prompt Engineering, upgrade ke CoddyKit PRO. Kursus AI Prompt Engineering mencakup 4 pelajaran total.

Apa yang akan aku pelajari di “Penjawaban Pertanyaan Visual”?

Ajukan pertanyaan spesifik tentang isi, jumlah, dan atribut gambar. Kamu berlatih AI Prompt Engineering dengan kode praktik yang langsung kamu jalankan di browser, dan tutor AI 24/7 menjawab pertanyaanmu saat kamu mengerjakan pelajaran ini.

Apakah aku perlu pengalaman untuk memulai AI Prompt Engineering?

Tidak diperlukan pengalaman sebelumnya. AI Prompt Engineering di CoddyKit dirancang untuk pemula hingga pelajar tingkat lanjut, jadi kamu bisa memulai di sini atau dari awal dan belajar sesuai kecepatan kamu sendiri. Ini adalah pelajaran 2 dari 4.

Berapa lama pelajaran “Penjawaban Pertanyaan Visual” memakan waktu?

Sebagian besar pelajaran CoddyKit memakan waktu sekitar 5–10 menit. Setiap pelajaran ringkas dan interaktif, jadi kamu membuat kemajuan stabil dan melanjutkan dari tempat kamu tinggalkan di web dan aplikasi.

Bisakah aku menulis dan menjalankan kode dalam pelajaran AI Prompt Engineering ini?

Ya. Setiap pelajaran AI Prompt Engineering menyertakan editor kode bawaan, jadi kamu menulis dan menjalankan kode nyata langsung di browser dan mendapatkan umpan balik AI instan — tidak diperlukan penyiapan lokal.

Semua pelajaran dalam kursus ini

  1. Perintah Deskripsi dan Pemberian Keterangan Gambar
  2. Penjawaban Pertanyaan Visual
  3. Perintah Perbandingan Banyak Gambar
  4. Perintah OCR dan Analisis Dokumen
← Kembali ke AI Prompt Engineering