Rate Limiting und Verwaltung von API-Kontingenten
Schützen Sie Agents in der Produktion vor Provider-Limits und unkontrollierten Kosten, indem Sie Anfragen drosseln, sie mit Backoff wiederholen und Kontingente pro Nutzer verwalten.
Rate Limiting und Verwaltung von API-Kontingenten ist eine kostenlose AI Agents with LangChain & Autonomous Workflows-Lektion auf CoddyKit. Dies ist Lektion 4 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des AI Agents with LangChain & Autonomous Workflows-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der AI Agents with LangChain & Autonomous Workflows-Kurs umfasst insgesamt 4 Lektionen.
Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.
The Limits You Face
LLM providers cap usage in two ways:
- Requests per minute (RPM)
- Tokens per minute (TPM)
Exceed them and calls return 429 Too Many Requests, breaking your agents under load.
Why Throttle Proactively
Waiting for 429s and retrying is wasteful. Proactive rate limiting spaces out requests so you stay under the cap, smoothing traffic and avoiding errors entirely.
Client-Side Rate Limiter
LangChain can throttle model calls with a built-in rate limiter that releases a fixed number of requests per second.
from langchain_core.rate_limiters import InMemoryRateLimiter
limiter = InMemoryRateLimiter(
requests_per_second=2,
max_bucket_size=5
)Attaching It to the Model
Pass the limiter to the chat model. Every call now waits its turn automatically.
llm = ChatOpenAI(
model='gpt-4o-mini',
rate_limiter=limiter
)Retry with Exponential Backoff
Some 429s and transient errors are unavoidable. Retry with growing delays so you do not hammer the provider.
llm_with_retry = llm.with_retry(
stop_after_attempt=5
)Respecting Retry-After
Providers often return a Retry-After header telling you how long to wait. Honoring it is more polite and effective than a fixed delay.
wait = int(response.headers.get('Retry-After', '1'))Per-User Quotas
Beyond provider limits, you set your own per-user quotas to control cost and fairness. Track usage in a store like Redis and reject or queue once a user exceeds their allowance.
used = redis.incr(f'quota:{user_id}')
if used > DAILY_LIMIT:
raise QuotaExceeded()Token Bucket Algorithm
The common pattern is a token bucket: tokens refill at a steady rate, each request consumes one, and an empty bucket means wait. It allows short bursts while enforcing an average rate.
Queuing Under Load
When demand spikes past your limits, queue requests instead of dropping them. A background worker drains the queue at a safe rate, keeping the system stable.
Spreading Across Keys
For high throughput you can rotate across multiple API keys or providers, distributing load so no single key hits its cap. Track each key's usage independently.
Monitoring Limits
Track 429 rates and how close you run to caps. Rising 429s signal you need a higher tier, better throttling, or more keys before users notice failures.
Quick Check
Test your rate-limiting knowledge.
Recap
You learned to manage limits and quotas in production:
- Providers cap RPM and TPM; 429s break agents
- Use an
InMemoryRateLimiterto throttle proactively - Add retry with backoff and honor
Retry-After - Enforce per-user quotas with a token bucket
- Queue, rotate keys, and monitor 429 rates
Good limit management keeps scaled agents reliable and affordable.
Lerne AI Agents with LangChain & Autonomous Workflows mit einem KI-Tutor — kostenlos
Schreibe und führe echten Code in deinem Browser aus, bekomme sofortige Hilfe von einem 24/7 KI-Tutor und setze dein Lernen im Web oder in der App fort.
- Kurse
- 12
- Lektionen
- 50
Häufig gestellte Fragen
Ist die Lektion „Rate Limiting und Verwaltung von API-Kontingenten“ kostenlos?
Ja — der vollständige Text von „Rate Limiting und Verwaltung von API-Kontingenten“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des AI Agents with LangChain & Autonomous Workflows-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der AI Agents with LangChain & Autonomous Workflows-Kurs umfasst insgesamt 4 Lektionen.
Was lerne ich in „Rate Limiting und Verwaltung von API-Kontingenten“?
Schützen Sie Agents in der Produktion vor Provider-Limits und unkontrollierten Kosten, indem Sie Anfragen drosseln, sie mit Backoff wiederholen und Kontingente pro Nutzer verwalten. Du übst AI Agents with LangChain & Autonomous Workflows mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.
Brauche ich Erfahrung, um AI Agents with LangChain & Autonomous Workflows zu starten?
Keine Vorkenntnisse erforderlich. AI Agents with LangChain & Autonomous Workflows auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 4 von 4.
Wie lange dauert die Lektion „Rate Limiting und Verwaltung von API-Kontingenten“?
Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.
Kann ich in dieser AI Agents with LangChain & Autonomous Workflows-Lektion Code schreiben und ausführen?
Ja. Jede AI Agents with LangChain & Autonomous Workflows-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.
Alle Lektionen in diesem Kurs
- Agenten auf Cloud-Plattformen deployen
- Agentenzustand und Sitzungen verwalten
- Agentenarchitekturen skalieren
- Rate Limiting und Verwaltung von API-Kontingenten