Alertas y SLO
Aprenda a convertir los datos de observabilidad en alertas accionables mediante objetivos de nivel de servicio, presupuestos de errores y alertas basadas en síntomas, reduciendo el ruido y avisando a las personas solo cuando sea importante.
Alertas y SLO es una lección gratuita de Microservices Communication Patterns (Saga, Circuit Breaker) en CoddyKit. Esta es la lección 4 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Microservices Communication Patterns (Saga, Circuit Breaker), y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Microservices Communication Patterns (Saga, Circuit Breaker) incluye 4 lecciones en total.
Partes de esta lección aún no han sido traducidas y se muestran en inglés.
From Observability to Action
Traces, logs, and metrics tell you what is happening. Alerting decides when a human needs to act. The goal is to page on real user pain, not on every blip.
What Is an SLI?
A Service Level Indicator is a measured aspect of service behavior, such as request latency, error rate, or availability. SLIs are the raw signal you alert on.
ok = 995
total = 1000
availability = ok / total * 100
print('Availability SLI:', availability, '%')What Is an SLO?
A Service Level Objective is a target for an SLI over a window, for example 'availability >= 99.9% over 30 days'. It defines what 'good enough' means.
Error Budgets
If your SLO is 99.9%, then 0.1% of requests are allowed to fail. That 0.1% is your error budget. Spend it on risk; if you burn it, slow down and stabilize.
slo = 99.9
budget_pct = 100 - slo
monthly_requests = 1000000
allowed_failures = monthly_requests * budget_pct / 100
print('Error budget (failures/month):', allowed_failures)Symptom vs Cause Alerts
Alert on symptoms users feel (high error rate, slow responses), not on every internal cause (one pod restarted). Cause alerts create noise; symptom alerts capture real impact.
Burn-Rate Alerting
Instead of paging the instant the SLO is missed, alert on how fast you are burning the error budget. A fast burn pages immediately; a slow burn opens a ticket.
def burn_rate(observed_error_rate, budget_rate):
return observed_error_rate / budget_rate
print('Burn rate:', burn_rate(0.01, 0.001), 'x budget')Reducing Alert Fatigue
Too many alerts and engineers ignore them all. Keep pages rare and meaningful:
- Every page must be actionable.
- Route non-urgent issues to tickets.
- Group related alerts together.
The Four Golden Signals
Google's SRE book recommends alerting on four signals:
- Latency
- Traffic
- Errors
- Saturation
Cover these and you catch most user-facing problems.
Severity Levels
Classify alerts by severity so the response matches the impact.
severity = {'P1': 'page on-call now', 'P2': 'page during hours', 'P3': 'create ticket'}
print(severity['P1'])Runbooks
Attach a runbook link to every alert. It tells the responder what the alert means, how to confirm impact, and the first steps to mitigate. Runbooks turn panic into procedure.
Closing the Loop
After each incident, review which alerts fired (and which should have). Tune thresholds, delete noisy alerts, and update runbooks. Alerting quality improves through iteration.
Quick Check
You have an SLO of 99.9% availability. What does the remaining 0.1% represent?
Recap
You learned alerting and SLOs:
- SLIs measure behavior; SLOs set targets; error budgets quantify allowed failure.
- Alert on symptoms and burn rate, not every cause.
- Use severity levels and runbooks to make pages actionable.
- Iterate to cut alert fatigue.
Good alerting pages humans only when users actually hurt.
Aprende Microservices Communication Patterns (Saga, Circuit Breaker) con un tutor de IA — gratis
Escribe y ejecuta código real en tu navegador, obtén ayuda instantánea de un tutor de IA disponible 24/7 y continúa donde lo dejaste en la web o en la aplicación.
- Cursos
- 12
- Lecciones
- 48
Preguntas frecuentes
¿La lección «Alertas y SLO» es gratis?
Sí — el texto completo de «Alertas y SLO» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Microservices Communication Patterns (Saga, Circuit Breaker), actualiza a CoddyKit PRO. El curso de Microservices Communication Patterns (Saga, Circuit Breaker) incluye 4 lecciones en total.
¿Qué aprenderé en «Alertas y SLO»?
Aprenda a convertir los datos de observabilidad en alertas accionables mediante objetivos de nivel de servicio, presupuestos de errores y alertas basadas en síntomas, reduciendo el ruido y avisando a… Practicas Microservices Communication Patterns (Saga, Circuit Breaker) con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.
¿Necesito experiencia previa para empezar Microservices Communication Patterns (Saga, Circuit Breaker)?
No se requiere experiencia previa. Microservices Communication Patterns (Saga, Circuit Breaker) en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 4 de 4.
¿Cuánto tiempo toma la lección «Alertas y SLO»?
La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.
¿Puedo escribir y ejecutar código en esta lección de Microservices Communication Patterns (Saga, Circuit Breaker)?
Sí. Cada lección de Microservices Communication Patterns (Saga, Circuit Breaker) incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.
Todas las lecciones de este curso
- Conceptos de trazabilidad distribuida
- Estrategias de registro centralizado
- Métricas y comprobaciones de estado
- Alertas y SLO