Health Checks & Server Monitoring
Implement health checks to automatically remove unhealthy backend servers from the load balancing pool.
Health Checks & Server Monitoring is a free API Gateway & Reverse Proxy (Nginx + Spring Cloud Gateway) lesson on CoddyKit — lesson 2 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the API Gateway & Reverse Proxy (Nginx + Spring Cloud Gateway) learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
Keeping Servers Healthy
Imagine a busy restaurant with multiple chefs. If one chef gets sick, you don't want to send orders to them, right?
In load balancing, health checks do exactly that. They constantly monitor backend servers to ensure they are ready to handle requests.
Why Health Checks Matter
Without health checks, a load balancer might keep sending requests to a server that is down, overloaded, or unresponsive.
- Bad User Experience: Users get errors instead of content.
- Wasted Resources: The load balancer tries to connect to a dead end.
- Cascading Failures: Overloads can spread if requests aren't redirected.
Nginx's Health Check Basics
Nginx provides built-in directives within its upstream block to perform passive health checks. These checks react to failed connection attempts or responses.
The two main directives are max_fails and fail_timeout.
`max_fails`: Failed Attempts
The max_fails directive sets the number of consecutive failed attempts after which Nginx considers a server "unhealthy".
- Default:
1(one failure marks it unhealthy). - `max_fails=0`: Disables health checks for that server.
- A "failure" can be a connection error, timeout, or specific HTTP status codes (like 5xx) if configured.
`fail_timeout`: Recovery Time
Once a server is marked unhealthy, Nginx will stop sending requests to it for a duration specified by fail_timeout.
- Default:
10s(10 seconds). - After this timeout, Nginx will tentatively send a request to the server to check if it has recovered.
- This prevents hammering a down server.
Implementing Nginx Health Checks
Let's see how to configure max_fails and fail_timeout in an Nginx upstream block. Here, we define two backend servers.
http {
upstream backend_servers {
server 192.168.1.100:8080 max_fails=3 fail_timeout=15s;
server 192.168.1.101:8080 max_fails=3 fail_timeout=15s;
}
server {
listen 80;
location / {
proxy_pass http://backend_servers;
}
}
}Nginx's Server Management
When a server hits its max_fails limit within the fail_timeout period, Nginx temporarily removes it from the load balancing pool.
- Requests are then distributed among the remaining healthy servers.
- After
fail_timeoutexpires, Nginx tries sending a single request to the "unhealthy" server. If it succeeds, the server is marked healthy again.
Observing Server Health
While Nginx's basic health checks are passive, you can observe their effects in Nginx logs.
Error logs will show messages when a server is marked down or up. For more advanced, active monitoring and a dashboard, Nginx Plus offers dedicated features, but that's beyond basic Nginx.
Health Check Tips
Properly configuring health checks is key for reliable systems:
- Tune Values: Adjust
max_failsandfail_timeoutbased on your application's responsiveness and recovery time. - Backend Readiness: Ensure your backend applications have a dedicated health endpoint (e.g.,
/health) that reports true service readiness. - Combine with Monitoring: Use external monitoring tools to alert you when Nginx marks servers down.
Quick Check on Health Checks
You've configured an Nginx upstream server with max_fails=2 and fail_timeout=30s. If the server fails 3 consecutive requests, what happens next?
Health Check Recap
In this lesson, we learned about Nginx's essential health check directives:
max_fails: The number of failed attempts before a server is marked unhealthy.fail_timeout: The period for which an unhealthy server is taken out of the load balancing pool.- These passive checks are crucial for maintaining backend reliability and improving user experience.
Frequently asked questions
Is the “Health Checks & Server Monitoring” lesson free?
Yes — the full text of “Health Checks & Server Monitoring” is free to read here on the web, and the API Gateway & Reverse Proxy (Nginx + Spring Cloud Gateway) course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the API Gateway & Reverse Proxy (Nginx + Spring Cloud Gateway) course, upgrade to CoddyKit PRO.
What will I learn in “Health Checks & Server Monitoring”?
Implement health checks to automatically remove unhealthy backend servers from the load balancing pool. You practise API Gateway & Reverse Proxy (Nginx + Spring Cloud Gateway) with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start API Gateway & Reverse Proxy (Nginx + Spring Cloud Gateway)?
No prior experience is required. API Gateway & Reverse Proxy (Nginx + Spring Cloud Gateway) on CoddyKit is structured for beginners through advanced learners; this is — lesson 2 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “Health Checks & Server Monitoring” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this API Gateway & Reverse Proxy (Nginx + Spring Cloud Gateway) lesson?
Yes. Every API Gateway & Reverse Proxy (Nginx + Spring Cloud Gateway) lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.
All lessons in this course
- Load Balancing Algorithms
- Health Checks & Server Monitoring
- Sticky Sessions & Session Persistence
- Weighted Load Balancing & Backup Servers