0Pricing
Claude Architect · Lesson

Severity Criteria with Examples

Anchor each severity level with a code example.

Severity Criteria with Examples is a free Claude Architect lesson on CoddyKit — lesson 3 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Claude Architect learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.

Why Severity Needs Criteria

When you ask Claude to review code, the weakest instruction you can give is a vague one: be more precise or only flag important issues. The model has no shared definition of "important", so its bar drifts from file to file.

The exam principle is blunt: explicit criteria beat vague adjectives. A severity scale (Critical / High / Medium / Low) is only useful if each level has a written rule the model can apply consistently — and the most reliable way to pin down a rule is to anchor it with a concrete code example.

This lesson builds a severity rubric for a CI/CD review agent, one level at a time, each level tied to an example.

The Failure Mode: Adjectives Without Anchors

Here is the kind of prompt that looks fine but performs badly. It names severity levels but never defines them, so the model guesses — and guesses differently each run.

The result is exactly the anti-pattern the exam warns about: noisy reviews where a missing null-check and a misspelled comment both get tagged "High". Reviewers stop trusting the labels.

system = (
    "You are a code reviewer. "
    "Rate each issue as Critical, High, Medium, or Low. "
    "Be precise and only report important problems."
)

# Problem: 'important', 'precise', and the four levels are
# never defined. The bar is whatever the model infers today.

The Fix: One Rule + One Example per Level

The repair is structural. For every severity level, give the model two things:

  • A rule — a testable condition ("causes data loss, security breach, or a crash in production").
  • An anchor example — a short snippet that unambiguously sits at that level.

This is few-shot prompting applied to a rubric: 2-4 targeted examples per ambiguity. The model generalizes from the anchors — it does not just echo them — so a handful of well-chosen examples calibrates the whole scale.

Critical — Anchor with a Security Example

Critical is reserved for issues that cause data loss, a security breach, or a production crash. Anchor it with something undeniable — here, raw string interpolation into SQL.

Notice the anchor does double duty: it defines the ceiling of the scale, so the model knows nothing milder should reach this level.

CRITICAL = """
Critical: causes data loss, a security breach, or a
production crash. Always report, even if low-confidence.

Example (SQL injection):
    query = f"SELECT * FROM users WHERE id = {user_input}"
    db.execute(query)
Why: user_input is interpolated unescaped -> injectable.
"""

High — Anchor with a Logic Bug

High covers wrong behavior that won't crash the process but produces incorrect results — a logic error, a broken edge case, an off-by-one. The anchor makes the boundary with Critical concrete: no breach, no crash, but the output is wrong.

HIGH = """
High: produces incorrect results or a test failure, but
does not breach security or crash production.

Example (off-by-one):
    for i in range(len(items) - 1):
        process(items[i])     # last item never processed
Why: range stops one element early; silent wrong output.
"""

Medium and Low — Anchor the Quiet End

The low end of the scale is where vague prompts leak the most false positives, so anchor it just as carefully.

  • Medium — maintainability or reliability risk that isn't yet a bug (a missing timeout, an unhandled-but-rare error path).
  • Low — style and naming only; no behavioral impact.

Defining Low explicitly is what lets you later say "don't report Low in pre-merge gates" without the model arguing.

MEDIUM = """
Medium: reliability or maintainability risk, not yet a bug.
Example:
    requests.get(url)        # no timeout -> can hang forever
"""

LOW = """
Low: style or naming only, no behavioral impact.
Example:
    def calc(x): return x*2  # name 'calc' is unclear
"""

Assemble the Rubric into the System Prompt

The anchored levels become one block in the system prompt. Keep this block stable and first — it's the same for every file you review, which makes it a perfect prompt-caching prefix. The per-file diff goes in the user turn, after the cached rubric.

import anthropic

client = anthropic.Anthropic()

system = [{
    "type": "text",
    "text": "You are a code reviewer.\n"
            + CRITICAL + HIGH + MEDIUM + LOW
            + "\nAssign exactly one level per finding using the\n"
              "rules and examples above. When unsure between two\n"
              "levels, pick the lower one.",
    "cache_control": {"type": "ephemeral"},
}]

Force Structure: Severity as an Enum

A written rubric tells the model how to decide; structured output guarantees the shape of the answer. Bind severity to a JSON Schema enum so the field can never be a free-text adjective like "prettyBad".

Exam rule to remember: mark a field required only if it is always present. severity and line always exist for a real finding, so they are required; an optional suggested_fix is not.

finding_schema = {
    "type": "object",
    "properties": {
        "line": {"type": "integer"},
        "severity": {
            "type": "string",
            "enum": ["critical", "high", "medium", "low"],
        },
        "rule": {"type": "string"},
        "suggested_fix": {"type": "string"},
    },
    "required": ["line", "severity", "rule"],
    "additionalProperties": False,
}

Wire the Rubric to the Review Call

Now combine the cached, anchored rubric with the enum-constrained schema in one request. The diff is the only volatile part, so it sits last in the user turn.

This pairing — explicit criteria for the decision, structured output for the format — is the exam's recommended pattern for reliable extraction and classification.

resp = client.messages.create(
    model="claude-opus-4-8",
    max_tokens=4096,
    thinking={"type": "adaptive"},
    system=system,                      # cached rubric prefix
    output_config={
        "format": {
            "type": "json_schema",
            "schema": {
                "type": "object",
                "properties": {"findings": {
                    "type": "array", "items": finding_schema}},
                "required": ["findings"],
                "additionalProperties": False,
            },
        }
    },
    messages=[{"role": "user", "content": diff_text}],
)

Severity Drives Gating, Not the Model

Once severity is a clean enum, the decision to block a merge is deterministic code, not a model judgment. The model classifies; your pipeline thresholds.

This mirrors the exam's hook principle: when a failure has real consequences (a broken merge), enforce it with deterministic code, not a probabilistic prompt. The rubric makes the model's labels trustworthy enough to gate on.

import json

findings = json.loads(resp.content[0].text)["findings"]

BLOCKING = {"critical", "high"}
blockers = [f for f in findings if f["severity"] in BLOCKING]

if blockers:
    print(f"BLOCK MERGE: {len(blockers)} issue(s)")
    raise SystemExit(1)
print("OK to merge (medium/low only)")

Don't Let the Model Self-Filter Severity

One subtle trap: telling the model at the finding stage to "only report Critical and High" suppresses recall — it silently drops issues it judged lower, and you lose coverage you might have wanted.

The robust pattern: have the model report every finding with its severity, then filter in a separate downstream step (your BLOCKING set, or an independent review pass). Coverage first, ranking second. Anchored criteria are what make that downstream ranking reliable.

Quick Check: Anchoring Severity

Apply the lesson to a realistic design choice.

Recap: Criteria You Can Point To

Key takeaways:

  • Vague adjectives drift; written rules don't. Replace "important" with a testable condition per severity level.
  • Anchor every level with a code example. 2-4 targeted few-shot anchors calibrate the scale — the model generalizes from them.
  • Lock the label with an enum. Structured output makes severity always one of the valid values; require only fields that are always present.
  • Cache the rubric, vary the diff. Keep the stable criteria first in the system prompt; put the per-file code last.
  • Classify in the model, gate in code. Report all findings with severity, then filter/block deterministically downstream — never self-filter at the finding stage.

Frequently asked questions

Is the “Severity Criteria with Examples” lesson free?

Yes — the full text of “Severity Criteria with Examples” is free to read here on the web, and the Claude Architect course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Claude Architect course, upgrade to CoddyKit PRO.

What will I learn in “Severity Criteria with Examples”?

Anchor each severity level with a code example. You practise Claude Architect with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.

Do I need any experience to start Claude Architect?

No prior experience is required. Claude Architect on CoddyKit is structured for beginners through advanced learners; this is — lesson 3 of 4, so you can start here or from the beginning and move at your own pace.

How long does the “Severity Criteria with Examples” lesson take?

Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.

Can I write and run code in this Claude Architect lesson?

Yes. Every Claude Architect lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.

All lessons in this course

  1. Explicit Criteria over Vague Instructions
  2. Categorical Examples
  3. Severity Criteria with Examples
  4. Reducing False Positives
← Back to Claude Architect