0Pricing
MLOps Academy · บทเรียน

ปรับขนาดอินสแตนซ์และเรพลิกาให้เหมาะสม

จับคู่ฮาร์ดแวร์กับรูปแบบภาระงานจริง

ปรับขนาดอินสแตนซ์และเรพลิกาให้เหมาะสม เป็นบทเรียน MLOps Academy ฟรีบน CoddyKit นี่คือบทเรียนที่ 1 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน MLOps Academy และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส MLOps Academy มีบทเรียนทั้งหมด 4 บทเรียน

บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ

Pay for What You Use

Most ML serving bills come from instances sitting half-idle. Right-sizing means matching hardware and replica count to your real load, not your worst-case fear.

Measure Before You Cut

You cannot right-size what you have not measured. Start by watching real CPU and memory utilization over a normal traffic day.

The Overprovisioning Trap

Picking a giant instance just in case feels safe but burns cash every hour. Chronically low utilization is the clearest sign of overprovisioning.

Read Your Utilization

If average CPU sits near 15 percent, your instance is far too big. Aim for a healthy target utilization with headroom for short spikes.

kubectl top pods -n serving

Vertical vs Horizontal

Vertical scaling gives one replica a bigger machine. Horizontal scaling adds more small replicas instead, which usually handles bursty traffic better.

Set Requests and Limits

On Kubernetes, your resource request tells the scheduler how much each replica truly needs. Set it from observed usage, not a round guess.

resources:
  requests:
    cpu: "500m"
    memory: "1Gi"

Pick the Right Replica Count

Too few replicas means queued requests and slow responses. Too many means idle pods you still pay for. Size replicas to your peak concurrent load.

Let Autoscaling Track Demand

A Horizontal Pod Autoscaler adds replicas when load rises and removes them when it falls, so you stop paying for off-peak capacity.

kubectl autoscale deploy model --min 2 --max 10 --cpu-percent 60

Keep a Safety Floor

A minimum replica count keeps a few pods warm so traffic never hits a cold, empty service. This is your trade-off between cost and availability.

GPU Boxes Are Pricey

GPU instances cost many times more than CPU. Only request a GPU when your model truly needs it, and pack work tightly so it never sits idle.

Right-Sizing Is Ongoing

Traffic patterns drift over weeks and months. Revisit instance type and replica counts on a schedule so your fleet stays a good fit.

Quick Check

Let us see what low utilization is telling you.

Recap

You measured utilization, chose between vertical and horizontal scaling, set requests, and added autoscaling. Your fleet now matches real demand. 💸

คำถามที่พบบ่อย

บทเรียน “ปรับขนาดอินสแตนซ์และเรพลิกาให้เหมาะสม” ฟรีหรือไม่

ใช่ — ข้อความเต็มของ “ปรับขนาดอินสแตนซ์และเรพลิกาให้เหมาะสม” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส MLOps Academy ให้อัปเกรดเป็น CoddyKit PRO คอร์ส MLOps Academy มีบทเรียนทั้งหมด 4 บทเรียน

คุณจะเรียนรู้อะไรในบทเรียน “ปรับขนาดอินสแตนซ์และเรพลิกาให้เหมาะสม”

จับคู่ฮาร์ดแวร์กับรูปแบบภาระงานจริง คุณปฏิบัติ MLOps Academy ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน

คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน MLOps Academy หรือไม่

ไม่จำเป็นต้องมีประสบการณ์มาก่อน MLOps Academy บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 1 จากทั้งหมด 4 บทเรียน

บทเรียน “ปรับขนาดอินสแตนซ์และเรพลิกาให้เหมาะสม” ใช้เวลานานแค่ไหน

บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย

ฉันเขียนและรันโค้ดในบทเรียน MLOps Academy นี้ได้ไหม

ได้ บทเรียน MLOps Academy ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ

บทเรียนทั้งหมดในหลักสูตรนี้

  1. ปรับขนาดอินสแตนซ์และเรพลิกาให้เหมาะสม
  2. ทำควอนไทซ์และกลั่นโมเดลเพื่อให้อนุมานถูกลง
  3. ใช้ Spot Instances สำหรับการฝึก
  4. ติดตามต้นทุนต่อผลพยากรณ์
← กลับไปที่ MLOps Academy