Instanzen und Replikas passend dimensionieren
Passen Sie die Hardware an reale Lastprofile an.
Instanzen und Replikas passend dimensionieren ist eine kostenlose MLOps Academy-Lektion auf CoddyKit. Dies ist Lektion 1 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des MLOps Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der MLOps Academy-Kurs umfasst insgesamt 4 Lektionen.
Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.
Pay for What You Use
Most ML serving bills come from instances sitting half-idle. Right-sizing means matching hardware and replica count to your real load, not your worst-case fear.
Measure Before You Cut
You cannot right-size what you have not measured. Start by watching real CPU and memory utilization over a normal traffic day.
The Overprovisioning Trap
Picking a giant instance just in case feels safe but burns cash every hour. Chronically low utilization is the clearest sign of overprovisioning.
Read Your Utilization
If average CPU sits near 15 percent, your instance is far too big. Aim for a healthy target utilization with headroom for short spikes.
kubectl top pods -n servingVertical vs Horizontal
Vertical scaling gives one replica a bigger machine. Horizontal scaling adds more small replicas instead, which usually handles bursty traffic better.
Set Requests and Limits
On Kubernetes, your resource request tells the scheduler how much each replica truly needs. Set it from observed usage, not a round guess.
resources:
requests:
cpu: "500m"
memory: "1Gi"Pick the Right Replica Count
Too few replicas means queued requests and slow responses. Too many means idle pods you still pay for. Size replicas to your peak concurrent load.
Let Autoscaling Track Demand
A Horizontal Pod Autoscaler adds replicas when load rises and removes them when it falls, so you stop paying for off-peak capacity.
kubectl autoscale deploy model --min 2 --max 10 --cpu-percent 60Keep a Safety Floor
A minimum replica count keeps a few pods warm so traffic never hits a cold, empty service. This is your trade-off between cost and availability.
GPU Boxes Are Pricey
GPU instances cost many times more than CPU. Only request a GPU when your model truly needs it, and pack work tightly so it never sits idle.
Right-Sizing Is Ongoing
Traffic patterns drift over weeks and months. Revisit instance type and replica counts on a schedule so your fleet stays a good fit.
Quick Check
Let us see what low utilization is telling you.
Recap
You measured utilization, chose between vertical and horizontal scaling, set requests, and added autoscaling. Your fleet now matches real demand. 💸
Häufig gestellte Fragen
Ist die Lektion „Instanzen und Replikas passend dimensionieren“ kostenlos?
Ja — der vollständige Text von „Instanzen und Replikas passend dimensionieren“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des MLOps Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der MLOps Academy-Kurs umfasst insgesamt 4 Lektionen.
Was lerne ich in „Instanzen und Replikas passend dimensionieren“?
Passen Sie die Hardware an reale Lastprofile an. Du übst MLOps Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.
Brauche ich Erfahrung, um MLOps Academy zu starten?
Keine Vorkenntnisse erforderlich. MLOps Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 1 von 4.
Wie lange dauert die Lektion „Instanzen und Replikas passend dimensionieren“?
Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.
Kann ich in dieser MLOps Academy-Lektion Code schreiben und ausführen?
Ja. Jede MLOps Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.
Alle Lektionen in diesem Kurs
- Instanzen und Replikas passend dimensionieren
- Für günstigere Inferenz quantisieren und destillieren
- Spot-Instanzen für das Training nutzen
- Kosten pro Vorhersage erfassen