Rate Limiting and Throttling
Learn how rate limiting protects systems from abuse and overload, including the token bucket and sliding window algorithms.
Rate Limiting and Throttling is a free System Design Basics for Backend Developers lesson on CoddyKit — lesson 4 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the System Design Basics for Backend Developers learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
Why Rate Limit?
Rate limiting caps how many requests a client can make in a time window. It protects a system from abuse, accidental floods, and runaway clients.
- Stops brute-force and scraping attacks
- Ensures fair sharing among clients
- Protects backends from overload
Rate Limiting vs Throttling
The terms overlap but differ slightly: rate limiting rejects requests over a hard cap, while throttling often slows or queues excess requests rather than rejecting them outright.
Fixed Window Counter
The simplest scheme counts requests in fixed time windows, e.g. 100 per minute. It is easy but has an edge problem: a client can send 100 at the end of one window and 100 at the start of the next — 200 in a few seconds.
limit = 100
window = '12:00:00-12:00:59'
count = 0
# reset count to 0 each new windowSliding Window
A sliding window smooths the edge problem by weighting the previous window or tracking timestamps over a rolling interval. It gives a more accurate, fairer limit at the cost of more bookkeeping.
Token Bucket
The token bucket is the most popular algorithm. Tokens refill at a steady rate up to a capacity. Each request consumes a token; if the bucket is empty, the request is rejected. This allows short bursts while bounding the average rate.
import time
class Bucket:
def __init__(self, cap, rate):
self.cap = cap
self.rate = rate
self.tokens = cap
self.last = time.time()
def allow(self):
now = time.time()
self.tokens = min(self.cap, self.tokens + (now - self.last) * self.rate)
self.last = now
if self.tokens >= 1:
self.tokens -= 1
return True
return False
b = Bucket(5, 1)
print([b.allow() for _ in range(7)])Leaky Bucket
The leaky bucket processes requests at a fixed rate, queuing bursts and 'leaking' them out steadily. It smooths traffic into a constant outflow — good when the downstream needs a steady, predictable load.
Choosing the Limit Key
Decide what to limit on:
- Per API key or user — fair per-account limits
- Per IP — defends against anonymous abuse
- Per endpoint — protects expensive operations
Often you combine several keys.
Communicating Limits
Tell clients their status with standard headers and the right status code, so well-behaved clients can back off.
HTTP/1.1 429 Too Many Requests
Retry-After: 30
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1735689600Distributed Rate Limiting
With many app servers, an in-memory counter per server is inconsistent. Use a shared store like Redis with atomic increments (or Lua scripts) so the limit is enforced globally across the fleet.
INCR rl:user:42
EXPIRE rl:user:42 60
# reject when value > limitRate Limiting and DDoS
Rate limiting complements DDoS protection. Application-layer limits stop a single abusive client, while edge and network defenses absorb large volumetric floods before they reach your servers. Defense in depth uses both.
Designing Good Limits
Set limits from real usage data, allow reasonable bursts, expose clear headers, and return 429 with Retry-After. Consider tiered limits — higher caps for paid plans, stricter ones for unauthenticated traffic.
Quick Check
Test your understanding of rate limiting.
Recap
You learned to protect systems with rate limiting:
- Fixed window, sliding window, token bucket, and leaky bucket
- Choose limit keys: per user, per IP, per endpoint
- Return 429 with Retry-After and rate-limit headers
- Use a shared store like Redis for distributed enforcement
Frequently asked questions
Is the “Rate Limiting and Throttling” lesson free?
Yes — the full text of “Rate Limiting and Throttling” is free to read here on the web, and the System Design Basics for Backend Developers course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the System Design Basics for Backend Developers course, upgrade to CoddyKit PRO.
What will I learn in “Rate Limiting and Throttling”?
Learn how rate limiting protects systems from abuse and overload, including the token bucket and sliding window algorithms. You practise System Design Basics for Backend Developers with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start System Design Basics for Backend Developers?
No prior experience is required. System Design Basics for Backend Developers on CoddyKit is structured for beginners through advanced learners; this is — lesson 4 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “Rate Limiting and Throttling” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this System Design Basics for Backend Developers lesson?
Yes. Every System Design Basics for Backend Developers lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.
All lessons in this course
- Authentication & Authorization
- Data Encryption & Privacy
- DDoS Protection & Firewalls
- Rate Limiting and Throttling