0Pricing
MLOps Academy · บทเรียน

ลดขนาดเป็นศูนย์และเพิ่มกลับขึ้นมา

ประหยัดต้นทุนด้วยการปรับขนาดอัตโนมัติตามคำขอ

ลดขนาดเป็นศูนย์และเพิ่มกลับขึ้นมา เป็นบทเรียน MLOps Academy ฟรีบน CoddyKit นี่คือบทเรียนที่ 2 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน MLOps Academy และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส MLOps Academy มีบทเรียนทั้งหมด 4 บทเรียน

บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ

Idle Models Cost Money

A model pod sitting with no traffic still burns CPU, memory, and cloud bills. KServe can shrink an idle service all the way down. Scale to zero stops that waste. 💸

Powered by Knative

KServe's serverless mode rides on Knative, which watches request volume and adjusts replicas. When traffic stops, it can remove every pod for that service.

Request-Driven Autoscaling

Replicas track demand, not a fixed schedule. As requests rise, KServe adds pods, and as they fade it scales back. This is request-driven autoscaling in action.

Setting minReplicas to Zero

To allow full scale down, you set minReplicas to 0 in the predictor spec. With zero allowed, an idle service drops to no running pods at all.

spec:
  predictor:
    minReplicas: 0
    model:
      modelFormat:
        name: sklearn

The Concurrency Target

The autoscaler aims for a target number of in-flight requests per pod. This concurrency target decides how aggressively KServe adds replicas under load.

metadata:
  annotations:
    autoscaling.knative.dev/target: "10"

Scaling Back Up

When a request arrives at a zero-scaled service, Knative spins a pod up on demand. Traffic resumes the moment the pod is ready, so the scale up is automatic.

The Cold Start Cost

That first request after zero must wait for the pod to start and load the model. This delay is the cold start, the main trade-off of scaling to zero. ⏱️

When Zero Makes Sense

Scale to zero shines for bursty or rare workloads where idle time dominates. For steady, latency-critical traffic, the cold start penalty may not be worth it.

Keep One Warm Instead

If cold starts hurt, set minReplicas to 1 so one pod always stays alive. You trade a little cost for a guaranteed warm instance ready to serve.

spec:
  predictor:
    minReplicas: 1

Cap the Upper End

You can also bound growth with maxReplicas so a traffic spike never overruns your cluster budget. It puts a ceiling on the autoscaler.

spec:
  predictor:
    minReplicas: 0
    maxReplicas: 5

Watch It Scale

You can confirm the behavior by watching pods appear and vanish with traffic. The pod count drops to zero when idle and climbs back when calls come in.

kubectl get pods -l serving.kserve.io/inferenceservice=sklearn-iris -w

Quick Check

What is the main downside of letting a service scale to zero?

Recap

You saw how minReplicas: 0 lets KServe scale idle models to zero and back on demand. Cap with maxReplicas, and keep one warm if cold starts hurt. Great progress! 🎉

คำถามที่พบบ่อย

บทเรียน “ลดขนาดเป็นศูนย์และเพิ่มกลับขึ้นมา” ฟรีหรือไม่

ใช่ — ข้อความเต็มของ “ลดขนาดเป็นศูนย์และเพิ่มกลับขึ้นมา” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส MLOps Academy ให้อัปเกรดเป็น CoddyKit PRO คอร์ส MLOps Academy มีบทเรียนทั้งหมด 4 บทเรียน

คุณจะเรียนรู้อะไรในบทเรียน “ลดขนาดเป็นศูนย์และเพิ่มกลับขึ้นมา”

ประหยัดต้นทุนด้วยการปรับขนาดอัตโนมัติตามคำขอ คุณปฏิบัติ MLOps Academy ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน

คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน MLOps Academy หรือไม่

ไม่จำเป็นต้องมีประสบการณ์มาก่อน MLOps Academy บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 2 จากทั้งหมด 4 บทเรียน

บทเรียน “ลดขนาดเป็นศูนย์และเพิ่มกลับขึ้นมา” ใช้เวลานานแค่ไหน

บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย

ฉันเขียนและรันโค้ดในบทเรียน MLOps Academy นี้ได้ไหม

ได้ บทเรียน MLOps Academy ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ

บทเรียนทั้งหมดในหลักสูตรนี้

  1. ทรัพยากร InferenceService
  2. ลดขนาดเป็นศูนย์และเพิ่มกลับขึ้นมา
  3. เขียนตัวพยากรณ์แบบกำหนดเอง
  4. KServe เทียบกับ Seldon Core
← กลับไปที่ MLOps Academy