Visual Question Answering
Ponga domande specifiche sul contenuto, sulle quantità e sugli attributi delle immagini
Visual Question Answering è una lezione AI Prompt Engineering gratuita su CoddyKit. Questa è la lezione 2 di 4. Puoi leggere la lezione completa qui gratuitamente — poi esercitati direttamente nel browser con un editor di codice integrato e un tutor IA disponibile 24/7. Fa parte del percorso di apprendimento AI Prompt Engineering, e i tuoi progressi si sincronizzano tra il web e l'app CoddyKit. Il corso AI Prompt Engineering include 4 lezioni in totale.
Visual Question Answering
Il Visual Question Answering (VQA) è l'attività di rispondere a domande in linguaggio naturale su un'immagine. A differenza della descrizione di immagini, che descrive ogni elemento, il VQA concentra il modello sulla risposta a una domanda specifica.
I prompt VQA sono precisi e diretti e spesso richiedono di contare, identificare, confrontare o ragionare sui contenuti visivi. La qualità del prompt determina se otterrà una risposta precisa e utile oppure una risposta generica e vaga.
Struttura di base di un prompt VQA
Un prompt VQA associa un'immagine a una domanda specifica. La chiave consiste nel formulare la domanda con sufficiente precisione da ottenere una risposta diretta e utilizzabile:
import anthropic, base64
client = anthropic.Anthropic(api_key='YOUR_API_KEY')
def ask_about_image(image_path, question, answer_format='Direct answer. No extra explanation.'):
with open(image_path, 'rb') as f:
img_b64 = base64.standard_b64encode(f.read()).decode('utf-8')
prompt = f'{question}\n\n{answer_format}'
r = client.messages.create(
model='claude-opus-4-5', max_tokens=150,
messages=[{'role': 'user', 'content': [
{'type': 'image', 'source': {'type': 'base64', 'media_type': 'image/jpeg', 'data': img_b64}},
{'type': 'text', 'text': prompt}
]}]
)
return r.content[0].text
# Example VQA calls (replace image.jpg with actual image)
print('VQA function defined. Ready for image questions.')Domande sul conteggio
Il conteggio è un'attività VQA comune. I prompt per il conteggio formulati con precisione producono risultati più accurati rispetto a quelli vaghi:
# Vague (bad):
vague_prompt = 'How many people are there?'
# Precise (good): specifies what counts and what does not
counting_prompt = '''
How many people are visible in this image?
Count only: people whose faces OR bodies are at least 50% visible.
Do NOT count: people who are heavily cropped, cut off at the edge, or only partially visible.
Return a single number.
'''
# Even more precise: handles partial visibility explicitly
precise_count = '''
Count the number of distinct individuals visible in this image.
If a person is partially obscured, count them if more than half their body is visible.
Return JSON: {"count": integer, "partially_visible": integer, "notes": "string or null"}
'''
print('Counting prompts: vague vs precise.')
print('Precise prompts define edge cases explicitly.')Identificazione di marchi e loghi
Identificare i loghi dei marchi nelle immagini è un'attività comune nell'analisi dei prodotti. Il prompt deve specificare che cosa cercare e in quale formato restituirlo:
logo_prompt = '''
Identify all visible brand logos, company names, and product labels in this image.
For each, note:
- Brand/company name
- Where it appears in the image (top-left, center, on a product, etc.)
- Confidence: high (clearly legible) | medium (partially visible) | low (partially obscured)
Return JSON: {"brands": [{"name": str, "location": str, "confidence": str}]}
If no logos are visible, return: {"brands": []}
'''
import anthropic, base64, json
client = anthropic.Anthropic(api_key='YOUR_API_KEY')
def identify_brands(image_path):
with open(image_path, 'rb') as f:
img_b64 = base64.standard_b64encode(f.read()).decode('utf-8')
r = client.messages.create(
model='claude-opus-4-5', max_tokens=200,
messages=[{'role': 'user', 'content': [
{'type': 'image', 'source': {'type': 'base64', 'media_type': 'image/jpeg', 'data': img_b64}},
{'type': 'text', 'text': logo_prompt}
]}]
)
return json.loads(r.content[0].text)
print('Brand identification function defined.')Riconoscimento delle emozioni e delle espressioni
Identificare le espressioni emotive nelle immagini richiede una progettazione attenta del prompt, che tenga conto dell'incertezza:
emotion_prompt = '''
Describe the emotional expression of the person in this image.
Assess:
- Primary emotion: (happy, sad, angry, surprised, fearful, disgusted, neutral, or other)
- Intensity: (low, moderate, high)
- Confidence: (high if expression is clear, medium if subtle, low if face is obscured or turned away)
- Evidence: which specific facial features support your assessment
Return JSON:
{
"primary_emotion": str,
"intensity": str,
"confidence": str,
"evidence": str,
"secondary_emotion": str or null
}
If no person or face is clearly visible, return: {"primary_emotion": null, "confidence": "none", "reason": str}
'''
print(emotion_prompt)Domande sulle relazioni spaziali
Le domande sulla posizione relativa degli oggetti richiedono un vocabolario spaziale esplicito nel prompt:
spatial_prompt = '''
Answer questions about the spatial relationships of objects in this image.
Use these spatial terms consistently:
- Position in frame: top-left, top-center, top-right, middle-left, center, middle-right, bottom-left, bottom-center, bottom-right
- Relative position: in front of, behind, to the left of, to the right of, above, below, overlapping
- Distance: in the foreground, in the midground, in the background
Question: {question}
Answer in one or two sentences using the spatial vocabulary above.
'''
# Example questions:
questions = [
'Where is the red cup relative to the laptop?',
'Is the plant in the foreground or background?',
'What object is to the left of the person?'
]
for q in questions:
print(spatial_prompt.replace('{question}', q)[:200])
print('---')Valutazione della qualità e delle condizioni
Valutare la qualità o le condizioni degli oggetti nelle immagini è utile per l'ispezione dei prodotti, la valutazione immobiliare e il controllo qualità:
condition_prompt = '''
Assess the condition of the main subject in this image.
Rate on these dimensions (1-5 scale, 5=excellent):
- Physical condition: (1=heavily damaged, 5=like new)
- Cleanliness: (1=very dirty, 5=spotless)
- Completeness: (1=major parts missing, 5=fully intact)
For each rating, provide one-sentence evidence.
Return JSON:
{
"physical_condition": {"score": int, "evidence": str},
"cleanliness": {"score": int, "evidence": str},
"completeness": {"score": int, "evidence": str},
"overall_grade": "excellent|good|fair|poor",
"recommendation": str
}
'''
print('Condition assessment prompt defined.')
print('Useful for: product inspection, real estate, equipment maintenance.')Domande VQA con risposta sì/no
Le domande binarie sì/no richiedono prompt che impediscano al modello di fornire una risposta in prosa piena di riserve quando serve un semplice valore booleano:
def yes_no_question(image_path, question):
with open(image_path, 'rb') as f:
img_b64 = base64.standard_b64encode(f.read()).decode('utf-8')
prompt = f'''
Answer this yes/no question about the image.
Return JSON: {{"answer": "yes|no", "confidence": "high|medium|low", "reason": str}}
Do NOT answer with maybe, possibly, or a hedged statement.
If you genuinely cannot determine the answer, return {{"answer": "unclear", "confidence": "low", "reason": str}}
Question: {question}
'''
r = client.messages.create(
model='claude-opus-4-5', max_tokens=100,
messages=[{'role': 'user', 'content': [
{'type': 'image', 'source': {'type': 'base64', 'media_type': 'image/jpeg', 'data': img_b64}},
{'type': 'text', 'text': prompt}
]}]
)
return json.loads(r.content[0].text)
# Example: 'Is there a safety helmet visible in the image?'
print('Yes/no VQA function defined.')Concatenare domande VQA
È possibile concatenare più domande VQA sulla stessa immagine in un unico prompt per ridurre il numero di chiamate API:
multi_question_prompt = '''
Answer all of the following questions about this image.
Return a JSON object where each key is the question ID.
Questions:
1. How many people are visible?
2. What is the approximate age range of the youngest person?
3. Is there any food visible in the image?
4. What is the dominant color in the image?
5. Is the setting indoors or outdoors?
Return JSON:
{
"q1": {"answer": str},
"q2": {"answer": str},
"q3": {"answer": "yes|no", "details": str or null},
"q4": {"answer": str},
"q5": {"answer": "indoors|outdoors|unclear"}
}
'''
print('Multi-question VQA prompt — answers 5 questions in one API call.')Gestire l'incertezza nel VQA
A volte non è possibile rispondere con certezza alle domande VQA: l'immagine potrebbe essere sfocata, l'elemento rilevante potrebbe essere parzialmente nascosto oppure la risposta potrebbe essere effettivamente ambigua. Chieda al modello di esprimere esplicitamente l'incertezza invece di obbligarlo a tirare a indovinare:
uncertainty_vqa_prompt = '''
Answer this question about the image as precisely as possible.
If the answer is not clearly visible or is ambiguous, say so explicitly.
Question: {question}
Return JSON:
{
"answer": str,
"confidence": "high|medium|low|cannot_determine",
"limitation": str or null
}
For confidence levels:
- high: Answer is clearly visible and unambiguous
- medium: Visible but some uncertainty
- low: Partially visible or requires inference
- cannot_determine: Not enough visual information
Question: What brand is printed on the water bottle?
'''
print(uncertainty_vqa_prompt)Prompt VQA specifici del dominio
Domini diversi richiedono un vocabolario VQA e standard di misurazione differenti. I prompt specifici del dominio producono risposte più accurate e utilizzabili:
# Manufacturing quality control VQA
qc_prompt = '''
Inspect this product image for quality defects.
Answer each question:
1. Are there any visible scratches or surface damage? (yes/no + location)
2. Is the product alignment within expected tolerance? (yes/no)
3. Are all required labels/markings present? (yes/no + list missing ones)
4. Overall QC result: PASS or FAIL?
Return JSON:
{"scratches": {"present": bool, "location": str or null},
"alignment_ok": bool,
"labels_complete": bool, "missing_labels": [str],
"qc_result": "PASS|FAIL",
"fail_reasons": [str]}
'''
# Food safety VQA
food_prompt = '''
Inspect this food preparation image.
1. Are gloves being worn? 2. Is hair covered? 3. Any visible contamination risk?
Return JSON: {"gloves": bool, "hair_covered": bool, "contamination_risk": bool, "details": str}
'''
print("Domain-specific QC and food safety VQA prompts defined.")Verifica rapida
Quale prompt VQA ha maggiori probabilità di produrre una risposta precisa e utilizzabile quando si contano oggetti in un'immagine?
Prompt VQA — Punti chiave
Per ottenere risultati efficaci nel Visual Question Answering, è necessario progettare i prompt con precisione:
- Formuli domande specifiche e dirette — eviti termini vaghi come alcuni o vari
- Definisca esplicitamente i casi limite nelle domande che richiedono un conteggio (che cosa conta come parzialmente visibile?)
- Specifichi il formato esatto dell'output — JSON, un singolo numero, sì/no — per evitare risposte discorsive e ambigue
- Includa il livello di confidenza per tutte le risposte, così da poter segnalare gli output incerti
- Raggruppi in un'unica chiamata le domande relative alla stessa immagine, per ridurre i costi dell'API
- Per le domande binarie con risposta sì/no, impedisca esplicitamente le risposte ambigue prevedendo un'opzione di riserva non chiaro
- Il VQA specifico per un dominio (medico, legale, prodotti) richiede l'uso nel prompt del vocabolario del dominio
Impara AI Prompt Engineering con un tutor IA — gratis
Scrivi ed esegui vero codice nel tuo browser, ricevi aiuto istantaneo da un tutor IA disponibile 24/7, e riprendi da dove hai lasciato sul web o nell'app.
- Corsi
- 53
- Lezioni
- 199
Domande Frequenti
La lezione «Visual Question Answering» è gratuita?
Sì — il testo completo di «Visual Question Answering» è gratuito qui sul web. Per esercitarvi in modo interattivo (un editor di codice integrato e un tutor IA 24/7) e sbloccare il resto del corso AI Prompt Engineering, passa a CoddyKit PRO. Il corso AI Prompt Engineering include 4 lezioni in totale.
Cosa imparerò in «Visual Question Answering»?
Ponga domande specifiche sul contenuto, sulle quantità e sugli attributi delle immagini Eserciti AI Prompt Engineering con codice pratico che esegui direttamente nel browser, e un tutor IA 24/7 risponde alle tue domande mentre lavori sulla lezione.
Ho bisogno di esperienza per iniziare AI Prompt Engineering?
Non è richiesta alcuna esperienza precedente. AI Prompt Engineering su CoddyKit è strutturato per principianti e studenti avanzati, quindi puoi iniziare da qui o dall'inizio e procedere al tuo ritmo. Questa è la lezione 2 di 4.
Quanto tempo richiede la lezione «Visual Question Answering»?
La maggior parte delle lezioni CoddyKit richiede circa 5–10 minuti. Ogni lezione è breve e interattiva, quindi fai progressi costanti e riprendi esattamente da dove hai lasciato su web e app.
Posso scrivere ed eseguire codice in questa lezione AI Prompt Engineering?
Sì. Ogni lezione AI Prompt Engineering include un editor di codice integrato, quindi scrivi ed esegui codice reale direttamente nel tuo browser e ricevi feedback istantaneo dall'IA — nessuna configurazione locale necessaria.
Tutte le lezioni di questo corso
- Prompt per descrivere e sottotitolare immagini
- Visual Question Answering
- Prompt per il confronto tra più immagini
- Prompt per OCR e analisi dei documenti