0Pricing
FastAPI Backend Development Bootcamp · Lekcja

Alerty dotyczące SLO i budżetów błędów

Definiuj cele poziomu usług i konfiguruj użyteczne alerty uruchamiane, zanim użytkownicy zauważą pogorszenie działania.

Alerty dotyczące SLO i budżetów błędów to bezpłatna lekcja FastAPI Backend Development Bootcamp na CoddyKit. To lekcja 4 z 4. Możesz przeczytać całą lekcję poniżej za darmo — a potem ćwiczyć ją interaktywnie w przeglądarce z wbudowanym edytorem kodu i tutorem AI dostępnym 24/7. To część ścieżki edukacyjnej FastAPI Backend Development Bootcamp, a Twój postęp synchronizuje się między webem a aplikacją CoddyKit. Kurs FastAPI Backend Development Bootcamp zawiera 4 lekcji w sumie.

Części tej lekcji nie zostały jeszcze przetłumaczone i są wyświetlane po angielsku.

Why Alert on SLOs, Not Raw Metrics

Traditional alerts fire on raw symptoms like CPU > 90% or error_count > 100. The problem: these page you for things users never notice, and stay silent during slow degradation that hurts customers.

SLO-based alerting inverts this. You first define what "good service" means to a user, then alert only when you are at risk of breaking that promise.

  • SLI (Service Level Indicator): a measured ratio, e.g. fraction of fast, successful requests.
  • SLO (Service Level Objective): the target for that SLI, e.g. 99.9% over 30 days.
  • Error budget: the allowed failure, i.e. 100% minus the SLO.

In this lesson you will define SLOs for a FastAPI service and wire alerts that fire before users notice degradation.

Picking a Good SLI for an API

A good SLI is a ratio of good events to valid events, scaled 0 to 100%. For a FastAPI backend the two workhorse SLIs are:

  • Availability: successful responses / all valid responses. Treat 5xx as failures; usually exclude 4xx (client's fault).
  • Latency: requests served under a threshold / all requests, e.g. responses faster than 300ms.

Below is a tiny, self-contained calculator that turns raw request logs into these two SLIs.

def compute_slis(requests, latency_threshold_ms=300):
    valid = [r for r in requests if r["status"] < 500 or r["status"] >= 500]
    total = len(valid)
    good_avail = sum(1 for r in valid if r["status"] < 500)
    fast = sum(1 for r in valid if r["latency_ms"] <= latency_threshold_ms)
    availability = good_avail / total
    latency_sli = fast / total
    return {"availability": availability, "latency": latency_sli}


sample = [
    {"status": 200, "latency_ms": 120},
    {"status": 200, "latency_ms": 410},
    {"status": 500, "latency_ms": 90},
    {"status": 200, "latency_ms": 250},
    {"status": 503, "latency_ms": 600},
]

slis = compute_slis(sample)
print(f"availability = {slis['availability']:.2%}")
print(f"latency      = {slis['latency']:.2%}")

From SLO to Error Budget

Once you commit to an SLO, the error budget is simply the failures you are allowed before breaching it.

  • SLO 99.9% over 30 days → budget = 0.1% of all requests may fail.
  • If you serve 10,000,000 requests/month, that is 10,000 failed requests you can spend.
  • Expressed as time at full outage: 30 days × 0.1% ≈ 43 minutes of downtime per month.

The budget is a currency. Every failure spends it. Alerting is about watching the burn rate of that currency.

def error_budget(slo_percent, total_requests, window_days=30):
    allowed_failure_ratio = 1 - (slo_percent / 100)
    budget_requests = total_requests * allowed_failure_ratio
    budget_minutes = window_days * 24 * 60 * allowed_failure_ratio
    return budget_requests, budget_minutes


for slo in (99.0, 99.9, 99.99):
    reqs, mins = error_budget(slo, 10_000_000)
    print(f"SLO {slo}% -> {reqs:,.0f} req budget, {mins:.1f} min downtime/30d")

Burn Rate: The Core Idea

Burn rate is how fast you are consuming the error budget relative to "on pace." A burn rate of 1 means you will exactly exhaust the 30-day budget in 30 days. A burn rate of 14.4 means you would exhaust it in roughly 2 days.

The formula:

  • burn_rate = observed_error_rate / (1 - SLO)
  • If SLO = 99.9%, the budget error rate is 0.001. An observed error rate of 1.44% gives a burn rate of 0.0144 / 0.001 = 14.4.

High burn rate over a short window = something is on fire right now. Modest burn over a long window = slow leak. You alert on both.

def burn_rate(observed_error_rate, slo_percent):
    budget_rate = 1 - (slo_percent / 100)
    return observed_error_rate / budget_rate


for err in (0.001, 0.005, 0.0144, 0.05):
    br = burn_rate(err, 99.9)
    hours_to_exhaust = (30 * 24) / br if br else float("inf")
    print(f"err={err:.3%} -> burn={br:5.1f}x, exhausts budget in {hours_to_exhaust:6.1f}h")

Multi-Window, Multi-Burn-Rate Alerts

A single threshold is a trap: alert too eagerly and you page on blips; too late and you miss real outages. The Google SRE-recommended pattern uses multiple burn rates over multiple windows:

  • Fast burn: 14.4x over 1 hour (and a 5-min check) → page now, you will burn the whole month's budget in ~2 days.
  • Slow burn: 6x over 6 hours → page, sustained leak.
  • Trickle: 1x over 3 days → ticket, not a page.

Pairing a long window with a short "is it still happening" window kills false positives: the alert only fires if the burn is both severe and current.

ALERT_TIERS = [
    {"name": "page_fast", "burn": 14.4, "long_window_h": 1, "short_window_m": 5},
    {"name": "page_slow", "burn": 6.0, "long_window_h": 6, "short_window_m": 30},
    {"name": "ticket",    "burn": 1.0, "long_window_h": 72, "short_window_m": 360},
]


def evaluate(long_burn, short_burn):
    for tier in ALERT_TIERS:
        if long_burn >= tier["burn"] and short_burn >= tier["burn"]:
            return tier["name"]
    return "ok"


print(evaluate(long_burn=16.0, short_burn=15.0))  # spike still active
print(evaluate(long_burn=16.0, short_burn=0.2))   # recovered, suppress
print(evaluate(long_burn=2.0,  short_burn=2.0))    # slow leak -> ticket

Instrumenting FastAPI to Feed the SLI

Alerts are only as good as the data. Add Prometheus counters and a histogram via middleware so every request contributes to your availability and latency SLIs.

This is framework code (FastAPI + middleware), so it is not standalone-runnable, but it is the canonical instrumentation pattern.

import time
from fastapi import FastAPI, Request
from prometheus_client import Counter, Histogram, make_asgi_app

app = FastAPI()

REQUESTS = Counter(
    "http_requests_total", "All requests", ["method", "path", "status"]
)
LATENCY = Histogram(
    "http_request_duration_seconds", "Request latency", ["path"],
    buckets=(0.05, 0.1, 0.3, 0.5, 1.0, 2.5),
)


@app.middleware("http")
async def record_metrics(request: Request, call_next):
    start = time.perf_counter()
    response = await call_next(request)
    elapsed = time.perf_counter() - start
    path = request.url.path
    REQUESTS.labels(request.method, path, response.status_code).inc()
    LATENCY.labels(path).observe(elapsed)
    return response


app.mount("/metrics", make_asgi_app())

Writing the SLI as a PromQL Recording Rule

With the counters in place, express the SLIs as PromQL. A recording rule precomputes the error ratio so alert rules stay cheap and readable.

The good-events ratio over the alert window is what you compare against the burn-rate threshold. Note how the threshold 0.0144 equals 14.4 × (1 - 0.999) — burn rate folded into the comparison.

# prometheus rules.yml
groups:
  - name: slo_recording
    rules:
      - record: job:http_error_ratio:rate1h
        expr: |
          sum(rate(http_requests_total{status=~"5.."}[1h]))
          /
          sum(rate(http_requests_total[1h]))

      - record: job:http_error_ratio:rate5m
        expr: |
          sum(rate(http_requests_total{status=~"5.."}[5m]))
          /
          sum(rate(http_requests_total[5m]))

The Fast-Burn Alert Rule

Now the alert. It fires only when both the 1h and 5m error ratios exceed the 14.4x burn threshold for a 99.9% SLO. The short window guarantees the problem is still happening before paging an on-call engineer.

  • 0.0144 = 14.4 × (1 - 0.999)
  • for: 2m debounces transient spikes.
  • Labels route severity; annotations carry the runbook link.
# prometheus alerts.yml
groups:
  - name: slo_alerts
    rules:
      - alert: HighErrorBudgetBurnFast
        expr: |
          job:http_error_ratio:rate1h > 0.0144
          and
          job:http_error_ratio:rate5m > 0.0144
        for: 2m
        labels:
          severity: page
          slo: availability-99.9
        annotations:
          summary: "Burning error budget 14.4x (1h+5m)"
          description: "At this rate the 30d budget is gone in ~2 days."
          runbook: "https://runbooks.internal/slo-fast-burn"

Simulating an Alert Decision in Python

Before trusting your rules in production, simulate them. This self-contained evaluator replays a window of requests, computes both short and long error ratios, and decides whether a page should fire — exactly mirroring the multi-window rule.

def error_ratio(window):
    if not window:
        return 0.0
    failures = sum(1 for r in window if r >= 500)
    return failures / len(window)


def should_page(long_window, short_window, slo=99.9, multiplier=14.4):
    threshold = multiplier * (1 - slo / 100)
    long_r = error_ratio(long_window)
    short_r = error_ratio(short_window)
    return long_r > threshold and short_r > threshold, long_r, short_r


long_w = [200] * 940 + [500] * 60   # 6% errors over the hour
short_w = [200] * 95 + [500] * 5     # 5% still in last 5 min

fire, lr, sr = should_page(long_w, short_w)
print(f"long={lr:.2%} short={sr:.2%} -> page={fire}")

Latency SLOs Need Percentiles, Not Averages

Averages hide tail pain. If 99% of requests take 80ms and 1% take 4s, the mean looks fine while real users rage. Define latency SLOs on percentiles (p95, p99) or as the fraction under a threshold.

With a Prometheus histogram, p99 over 5m is:

  • histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le))

Below, a pure-Python p-quantile helper shows the math your histogram approximates.

def percentile(values, p):
    if not values:
        return None
    s = sorted(values)
    k = (len(s) - 1) * (p / 100)
    lo = int(k)
    hi = min(lo + 1, len(s) - 1)
    frac = k - lo
    return s[lo] + (s[hi] - s[lo]) * frac


latencies = [70, 75, 80, 82, 90, 95, 110, 130, 300, 4000]
for p in (50, 95, 99):
    print(f"p{p} = {percentile(latencies, p):.0f} ms")

Actionable Alerts and Budget Policy

An alert that no one can act on is noise. Make every page actionable:

  • Symptom-based: alert on the SLO breach (users hurting), not on a single cause.
  • Runbook attached: each alert links steps to diagnose and mitigate.
  • Routed by severity: page wakes someone; ticket waits for business hours.

Tie it to an error-budget policy: when the budget is healthy, ship features fast. When it is exhausted, freeze risky deploys and prioritize reliability. The budget turns reliability from opinion into a shared, data-driven decision.

def budget_policy(consumed_ratio):
    if consumed_ratio < 0.5:
        return "GREEN: ship freely"
    if consumed_ratio < 0.9:
        return "YELLOW: extra review on risky changes"
    if consumed_ratio < 1.0:
        return "ORANGE: freeze non-critical deploys"
    return "RED: budget exhausted, reliability work only"


for used in (0.2, 0.7, 0.95, 1.1):
    print(f"{used:.0%} budget used -> {budget_policy(used)}")

Quick Check: Choosing the Alert Trigger

You run a FastAPI service with a 99.9% availability SLO. You want to page on-call only for problems that genuinely threaten the monthly error budget, while avoiding false alarms from brief blips. Which alerting strategy best fits?

Recap: SLO-Driven Alerting

You learned to alert on user-facing reliability instead of raw symptoms:

  • SLI = good events / valid events; SLO = the target; error budget = 100% minus the SLO.
  • Burn rate = observed error rate / budget rate; it tells you how fast the budget is being spent.
  • Multi-window, multi-burn-rate alerts (e.g. 14.4x over 1h+5m to page, slower burns to ticket) catch real degradation early without paging on blips.
  • Instrument FastAPI with Prometheus counters and histograms, express SLIs as recording rules, and define alerts on the error ratio — using percentiles, never averages, for latency.
  • Every alert must be actionable: symptom-based, runbook-linked, severity-routed, and backed by an error-budget policy that governs when to ship versus freeze.

Done well, your pages fire before users notice degradation — and stay quiet when nothing is wrong.

Często zadawane pytania

Czy lekcja „Alerty dotyczące SLO i budżetów błędów” jest bezpłatna?

Tak — pełny tekst „Alerty dotyczące SLO i budżetów błędów” jest dostępny za darmo tutaj w sieci. Aby ćwiczyć ją interaktywnie (wbudowany edytor kodu i tutor AI dostępny 24/7) i odblokować resztę kursu FastAPI Backend Development Bootcamp, przejdź na CoddyKit PRO. Kurs FastAPI Backend Development Bootcamp zawiera 4 lekcji w sumie.

Co nauczysz się w „Alerty dotyczące SLO i budżetów błędów”?

Definiuj cele poziomu usług i konfiguruj użyteczne alerty uruchamiane, zanim użytkownicy zauważą pogorszenie działania. Ćwiczysz FastAPI Backend Development Bootcamp z praktycznym kodem, który uruchamiasz bezpośrednio w przeglądarce, a tutor AI dostępny 24/7 odpowiada na Twoje pytania podczas pracy nad lekcją.

Czy potrzebuję doświadczenia, aby zacząć FastAPI Backend Development Bootcamp?

Nie wymagamy żadnego doświadczenia. FastAPI Backend Development Bootcamp w CoddyKit jest strukturyzowany dla początkujących i zaawansowanych użytkowników, więc możesz zacząć tutaj lub od początku i uczyć się w swoim tempie. To lekcja 4 z 4.

Ile czasu zajmuje lekcja „Alerty dotyczące SLO i budżetów błędów”?

Większość lekcji CoddyKit trwa około 5–10 minut. Każda lekcja to mały, interaktywny krok, dzięki czemu robisz systematyczne postępy i zawsze wracasz dokładnie do tego samego miejsca — na webie i w aplikacji.

Czy mogę pisać i uruchamiać kod w tej lekcji FastAPI Backend Development Bootcamp?

Tak. Każda lekcja FastAPI Backend Development Bootcamp zawiera wbudowany edytor kodu, więc piszesz i uruchamiasz prawdziwy kod bezpośrednio w przeglądarce i od razu otrzymujesz sprzężenie zwrotne od AI — bez konfiguracji na komputerze.

Wszystkie lekcje w tym kursie

  1. Ustrukturyzowane logowanie JSON i identyfikatory korelacji
  2. Śledzenie rozproszone za pomocą OpenTelemetry
  3. Metryki Prometheus i pulpity RED/USE
  4. Alerty dotyczące SLO i budżetów błędów
← Powrót do FastAPI Backend Development Bootcamp