Microservices Communication Patterns (Saga, Circuit Breaker) · บทเรียน

การแจ้งเตือนและ SLO

เรียนรู้วิธีเปลี่ยนข้อมูลด้านการสังเกตระบบให้เป็นการแจ้งเตือนที่นำไปดำเนินการได้ โดยใช้ Service Level Objectives งบประมาณข้อผิดพลาด และการแจ้งเตือนตามอาการ เพื่อลดสัญญาณรบกวนและเรียกคนเข้ามาดูแลเฉพาะเมื่อจำเป็น

บทเรียน 4 จาก 413 ขั้นตอน

การแจ้งเตือนและ SLO เป็นบทเรียน Microservices Communication Patterns (Saga, Circuit Breaker) ฟรีบน CoddyKit นี่คือบทเรียนที่ 4 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน Microservices Communication Patterns (Saga, Circuit Breaker) และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส Microservices Communication Patterns (Saga, Circuit Breaker) มีบทเรียนทั้งหมด 4 บทเรียน

บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ

From Observability to Action

Traces, logs, and metrics tell you what is happening. Alerting decides when a human needs to act. The goal is to page on real user pain, not on every blip.

What Is an SLI?

A Service Level Indicator is a measured aspect of service behavior, such as request latency, error rate, or availability. SLIs are the raw signal you alert on.

ok = 995
total = 1000
availability = ok / total * 100
print('Availability SLI:', availability, '%')

What Is an SLO?

A Service Level Objective is a target for an SLI over a window, for example 'availability >= 99.9% over 30 days'. It defines what 'good enough' means.

Error Budgets

If your SLO is 99.9%, then 0.1% of requests are allowed to fail. That 0.1% is your error budget. Spend it on risk; if you burn it, slow down and stabilize.

slo = 99.9
budget_pct = 100 - slo
monthly_requests = 1000000
allowed_failures = monthly_requests * budget_pct / 100
print('Error budget (failures/month):', allowed_failures)

Symptom vs Cause Alerts

Alert on symptoms users feel (high error rate, slow responses), not on every internal cause (one pod restarted). Cause alerts create noise; symptom alerts capture real impact.

Burn-Rate Alerting

Instead of paging the instant the SLO is missed, alert on how fast you are burning the error budget. A fast burn pages immediately; a slow burn opens a ticket.

def burn_rate(observed_error_rate, budget_rate):
    return observed_error_rate / budget_rate

print('Burn rate:', burn_rate(0.01, 0.001), 'x budget')

Reducing Alert Fatigue

Too many alerts and engineers ignore them all. Keep pages rare and meaningful:

  • Every page must be actionable.
  • Route non-urgent issues to tickets.
  • Group related alerts together.

The Four Golden Signals

Google's SRE book recommends alerting on four signals:

  • Latency
  • Traffic
  • Errors
  • Saturation

Cover these and you catch most user-facing problems.

Severity Levels

Classify alerts by severity so the response matches the impact.

severity = {'P1': 'page on-call now', 'P2': 'page during hours', 'P3': 'create ticket'}
print(severity['P1'])

Runbooks

Attach a runbook link to every alert. It tells the responder what the alert means, how to confirm impact, and the first steps to mitigate. Runbooks turn panic into procedure.

Closing the Loop

After each incident, review which alerts fired (and which should have). Tune thresholds, delete noisy alerts, and update runbooks. Alerting quality improves through iteration.

Quick Check

You have an SLO of 99.9% availability. What does the remaining 0.1% represent?

Recap

You learned alerting and SLOs:

  • SLIs measure behavior; SLOs set targets; error budgets quantify allowed failure.
  • Alert on symptoms and burn rate, not every cause.
  • Use severity levels and runbooks to make pages actionable.
  • Iterate to cut alert fatigue.

Good alerting pages humans only when users actually hurt.

เริ่มต้นได้ฟรี

เรียนรู้ Microservices Communication Patterns (Saga, Circuit Breaker) ด้วย AI tutor — ฟรี

เขียนและเรียกใช้โค้ดจริงในเบราว์เซอร์ของคุณ รับความช่วยเหลือทันทีจาก AI tutor 24/7 และเรียนรู้ต่อจากที่คุณหยุดบนเว็บหรือในแอป

คอร์ส
12
บทเรียน
48

คำถามที่พบบ่อย

บทเรียน “การแจ้งเตือนและ SLO” ฟรีหรือไม่

ใช่ — ข้อความเต็มของ “การแจ้งเตือนและ SLO” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส Microservices Communication Patterns (Saga, Circuit Breaker) ให้อัปเกรดเป็น CoddyKit PRO คอร์ส Microservices Communication Patterns (Saga, Circuit Breaker) มีบทเรียนทั้งหมด 4 บทเรียน

คุณจะเรียนรู้อะไรในบทเรียน “การแจ้งเตือนและ SLO”

เรียนรู้วิธีเปลี่ยนข้อมูลด้านการสังเกตระบบให้เป็นการแจ้งเตือนที่นำไปดำเนินการได้ โดยใช้ Service Level Objectives งบประมาณข้อผิดพลาด และการแจ้งเตือนตามอาการ เพื่อลดสัญญาณรบกวนและเรียกคนเข้ามาดูแลเฉพาะ… คุณปฏิบัติ Microservices Communication Patterns (Saga, Circuit Breaker) ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน

คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน Microservices Communication Patterns (Saga, Circuit Breaker) หรือไม่

ไม่จำเป็นต้องมีประสบการณ์มาก่อน Microservices Communication Patterns (Saga, Circuit Breaker) บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 4 จากทั้งหมด 4 บทเรียน

บทเรียน “การแจ้งเตือนและ SLO” ใช้เวลานานแค่ไหน

บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย

ฉันเขียนและรันโค้ดในบทเรียน Microservices Communication Patterns (Saga, Circuit Breaker) นี้ได้ไหม

ได้ บทเรียน Microservices Communication Patterns (Saga, Circuit Breaker) ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ

บทเรียนทั้งหมดในหลักสูตรนี้

  1. แนวคิดการติดตามแบบกระจาย
  2. กลยุทธ์การบันทึกข้อมูลแบบรวมศูนย์
  3. เมตริกและการตรวจสอบสถานะ
  4. การแจ้งเตือนและ SLO
← กลับไปที่ Microservices Communication Patterns (Saga, Circuit Breaker)