0Pricing
System Design Basics for Backend Developers · Lesson

Rate Limiting and Throttling

Learn how rate limiting protects systems from abuse and overload, including the token bucket and sliding window algorithms.

Rate Limiting and Throttling is a free System Design Basics for Backend Developers lesson on CoddyKit — lesson 4 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the System Design Basics for Backend Developers learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.

Why Rate Limit?

Rate limiting caps how many requests a client can make in a time window. It protects a system from abuse, accidental floods, and runaway clients.

  • Stops brute-force and scraping attacks
  • Ensures fair sharing among clients
  • Protects backends from overload

Rate Limiting vs Throttling

The terms overlap but differ slightly: rate limiting rejects requests over a hard cap, while throttling often slows or queues excess requests rather than rejecting them outright.

Fixed Window Counter

The simplest scheme counts requests in fixed time windows, e.g. 100 per minute. It is easy but has an edge problem: a client can send 100 at the end of one window and 100 at the start of the next — 200 in a few seconds.

limit = 100
window = '12:00:00-12:00:59'
count = 0
# reset count to 0 each new window

Sliding Window

A sliding window smooths the edge problem by weighting the previous window or tracking timestamps over a rolling interval. It gives a more accurate, fairer limit at the cost of more bookkeeping.

Token Bucket

The token bucket is the most popular algorithm. Tokens refill at a steady rate up to a capacity. Each request consumes a token; if the bucket is empty, the request is rejected. This allows short bursts while bounding the average rate.

import time
class Bucket:
    def __init__(self, cap, rate):
        self.cap = cap
        self.rate = rate
        self.tokens = cap
        self.last = time.time()
    def allow(self):
        now = time.time()
        self.tokens = min(self.cap, self.tokens + (now - self.last) * self.rate)
        self.last = now
        if self.tokens >= 1:
            self.tokens -= 1
            return True
        return False

b = Bucket(5, 1)
print([b.allow() for _ in range(7)])

Leaky Bucket

The leaky bucket processes requests at a fixed rate, queuing bursts and 'leaking' them out steadily. It smooths traffic into a constant outflow — good when the downstream needs a steady, predictable load.

Choosing the Limit Key

Decide what to limit on:

  • Per API key or user — fair per-account limits
  • Per IP — defends against anonymous abuse
  • Per endpoint — protects expensive operations

Often you combine several keys.

Communicating Limits

Tell clients their status with standard headers and the right status code, so well-behaved clients can back off.

HTTP/1.1 429 Too Many Requests
Retry-After: 30
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1735689600

Distributed Rate Limiting

With many app servers, an in-memory counter per server is inconsistent. Use a shared store like Redis with atomic increments (or Lua scripts) so the limit is enforced globally across the fleet.

INCR rl:user:42
EXPIRE rl:user:42 60
# reject when value > limit

Rate Limiting and DDoS

Rate limiting complements DDoS protection. Application-layer limits stop a single abusive client, while edge and network defenses absorb large volumetric floods before they reach your servers. Defense in depth uses both.

Designing Good Limits

Set limits from real usage data, allow reasonable bursts, expose clear headers, and return 429 with Retry-After. Consider tiered limits — higher caps for paid plans, stricter ones for unauthenticated traffic.

Quick Check

Test your understanding of rate limiting.

Recap

You learned to protect systems with rate limiting:

  • Fixed window, sliding window, token bucket, and leaky bucket
  • Choose limit keys: per user, per IP, per endpoint
  • Return 429 with Retry-After and rate-limit headers
  • Use a shared store like Redis for distributed enforcement

Frequently asked questions

Is the “Rate Limiting and Throttling” lesson free?

Yes — the full text of “Rate Limiting and Throttling” is free to read here on the web, and the System Design Basics for Backend Developers course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the System Design Basics for Backend Developers course, upgrade to CoddyKit PRO.

What will I learn in “Rate Limiting and Throttling”?

Learn how rate limiting protects systems from abuse and overload, including the token bucket and sliding window algorithms. You practise System Design Basics for Backend Developers with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.

Do I need any experience to start System Design Basics for Backend Developers?

No prior experience is required. System Design Basics for Backend Developers on CoddyKit is structured for beginners through advanced learners; this is — lesson 4 of 4, so you can start here or from the beginning and move at your own pace.

How long does the “Rate Limiting and Throttling” lesson take?

Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.

Can I write and run code in this System Design Basics for Backend Developers lesson?

Yes. Every System Design Basics for Backend Developers lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.

All lessons in this course

  1. Authentication & Authorization
  2. Data Encryption & Privacy
  3. DDoS Protection & Firewalls
  4. Rate Limiting and Throttling
← Back to System Design Basics for Backend Developers