Severity Criteria with Examples
Anchor each severity level with a code example.
Severity Criteria with Examples is a free Claude Architect lesson on CoddyKit — lesson 3 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Claude Architect learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
Why Severity Needs Criteria
When you ask Claude to review code, the weakest instruction you can give is a vague one: be more precise or only flag important issues. The model has no shared definition of "important", so its bar drifts from file to file.
The exam principle is blunt: explicit criteria beat vague adjectives. A severity scale (Critical / High / Medium / Low) is only useful if each level has a written rule the model can apply consistently — and the most reliable way to pin down a rule is to anchor it with a concrete code example.
This lesson builds a severity rubric for a CI/CD review agent, one level at a time, each level tied to an example.
The Failure Mode: Adjectives Without Anchors
Here is the kind of prompt that looks fine but performs badly. It names severity levels but never defines them, so the model guesses — and guesses differently each run.
The result is exactly the anti-pattern the exam warns about: noisy reviews where a missing null-check and a misspelled comment both get tagged "High". Reviewers stop trusting the labels.
system = (
"You are a code reviewer. "
"Rate each issue as Critical, High, Medium, or Low. "
"Be precise and only report important problems."
)
# Problem: 'important', 'precise', and the four levels are
# never defined. The bar is whatever the model infers today.The Fix: One Rule + One Example per Level
The repair is structural. For every severity level, give the model two things:
- A rule — a testable condition ("causes data loss, security breach, or a crash in production").
- An anchor example — a short snippet that unambiguously sits at that level.
This is few-shot prompting applied to a rubric: 2-4 targeted examples per ambiguity. The model generalizes from the anchors — it does not just echo them — so a handful of well-chosen examples calibrates the whole scale.
Critical — Anchor with a Security Example
Critical is reserved for issues that cause data loss, a security breach, or a production crash. Anchor it with something undeniable — here, raw string interpolation into SQL.
Notice the anchor does double duty: it defines the ceiling of the scale, so the model knows nothing milder should reach this level.
CRITICAL = """
Critical: causes data loss, a security breach, or a
production crash. Always report, even if low-confidence.
Example (SQL injection):
query = f"SELECT * FROM users WHERE id = {user_input}"
db.execute(query)
Why: user_input is interpolated unescaped -> injectable.
"""High — Anchor with a Logic Bug
High covers wrong behavior that won't crash the process but produces incorrect results — a logic error, a broken edge case, an off-by-one. The anchor makes the boundary with Critical concrete: no breach, no crash, but the output is wrong.
HIGH = """
High: produces incorrect results or a test failure, but
does not breach security or crash production.
Example (off-by-one):
for i in range(len(items) - 1):
process(items[i]) # last item never processed
Why: range stops one element early; silent wrong output.
"""Medium and Low — Anchor the Quiet End
The low end of the scale is where vague prompts leak the most false positives, so anchor it just as carefully.
- Medium — maintainability or reliability risk that isn't yet a bug (a missing timeout, an unhandled-but-rare error path).
- Low — style and naming only; no behavioral impact.
Defining Low explicitly is what lets you later say "don't report Low in pre-merge gates" without the model arguing.
MEDIUM = """
Medium: reliability or maintainability risk, not yet a bug.
Example:
requests.get(url) # no timeout -> can hang forever
"""
LOW = """
Low: style or naming only, no behavioral impact.
Example:
def calc(x): return x*2 # name 'calc' is unclear
"""Assemble the Rubric into the System Prompt
The anchored levels become one block in the system prompt. Keep this block stable and first — it's the same for every file you review, which makes it a perfect prompt-caching prefix. The per-file diff goes in the user turn, after the cached rubric.
import anthropic
client = anthropic.Anthropic()
system = [{
"type": "text",
"text": "You are a code reviewer.\n"
+ CRITICAL + HIGH + MEDIUM + LOW
+ "\nAssign exactly one level per finding using the\n"
"rules and examples above. When unsure between two\n"
"levels, pick the lower one.",
"cache_control": {"type": "ephemeral"},
}]Force Structure: Severity as an Enum
A written rubric tells the model how to decide; structured output guarantees the shape of the answer. Bind severity to a JSON Schema enum so the field can never be a free-text adjective like "prettyBad".
Exam rule to remember: mark a field required only if it is always present. severity and line always exist for a real finding, so they are required; an optional suggested_fix is not.
finding_schema = {
"type": "object",
"properties": {
"line": {"type": "integer"},
"severity": {
"type": "string",
"enum": ["critical", "high", "medium", "low"],
},
"rule": {"type": "string"},
"suggested_fix": {"type": "string"},
},
"required": ["line", "severity", "rule"],
"additionalProperties": False,
}Wire the Rubric to the Review Call
Now combine the cached, anchored rubric with the enum-constrained schema in one request. The diff is the only volatile part, so it sits last in the user turn.
This pairing — explicit criteria for the decision, structured output for the format — is the exam's recommended pattern for reliable extraction and classification.
resp = client.messages.create(
model="claude-opus-4-8",
max_tokens=4096,
thinking={"type": "adaptive"},
system=system, # cached rubric prefix
output_config={
"format": {
"type": "json_schema",
"schema": {
"type": "object",
"properties": {"findings": {
"type": "array", "items": finding_schema}},
"required": ["findings"],
"additionalProperties": False,
},
}
},
messages=[{"role": "user", "content": diff_text}],
)Severity Drives Gating, Not the Model
Once severity is a clean enum, the decision to block a merge is deterministic code, not a model judgment. The model classifies; your pipeline thresholds.
This mirrors the exam's hook principle: when a failure has real consequences (a broken merge), enforce it with deterministic code, not a probabilistic prompt. The rubric makes the model's labels trustworthy enough to gate on.
import json
findings = json.loads(resp.content[0].text)["findings"]
BLOCKING = {"critical", "high"}
blockers = [f for f in findings if f["severity"] in BLOCKING]
if blockers:
print(f"BLOCK MERGE: {len(blockers)} issue(s)")
raise SystemExit(1)
print("OK to merge (medium/low only)")Don't Let the Model Self-Filter Severity
One subtle trap: telling the model at the finding stage to "only report Critical and High" suppresses recall — it silently drops issues it judged lower, and you lose coverage you might have wanted.
The robust pattern: have the model report every finding with its severity, then filter in a separate downstream step (your BLOCKING set, or an independent review pass). Coverage first, ranking second. Anchored criteria are what make that downstream ranking reliable.
Quick Check: Anchoring Severity
Apply the lesson to a realistic design choice.
Recap: Criteria You Can Point To
Key takeaways:
- Vague adjectives drift; written rules don't. Replace "important" with a testable condition per severity level.
- Anchor every level with a code example. 2-4 targeted few-shot anchors calibrate the scale — the model generalizes from them.
- Lock the label with an enum. Structured output makes
severityalways one of the valid values; require only fields that are always present. - Cache the rubric, vary the diff. Keep the stable criteria first in the system prompt; put the per-file code last.
- Classify in the model, gate in code. Report all findings with severity, then filter/block deterministically downstream — never self-filter at the finding stage.
Frequently asked questions
Is the “Severity Criteria with Examples” lesson free?
Yes — the full text of “Severity Criteria with Examples” is free to read here on the web, and the Claude Architect course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Claude Architect course, upgrade to CoddyKit PRO.
What will I learn in “Severity Criteria with Examples”?
Anchor each severity level with a code example. You practise Claude Architect with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start Claude Architect?
No prior experience is required. Claude Architect on CoddyKit is structured for beginners through advanced learners; this is — lesson 3 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “Severity Criteria with Examples” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this Claude Architect lesson?
Yes. Every Claude Architect lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.
All lessons in this course
- Explicit Criteria over Vague Instructions
- Categorical Examples
- Severity Criteria with Examples
- Reducing False Positives