تشغيل التدريب كـ Kubernetes Job
نفّذ التدريب الدفعي حتى اكتماله على العنقود.
تشغيل التدريب كـ Kubernetes Job درس مجاني في MLOps Academy على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في MLOps Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة MLOps Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
Training Is Not a Server
A model server runs forever, but a training run should start, finish, and stop. Kubernetes has a different object built for that: the Job. 🏁
A Job Runs to Completion
A Job creates one or more Pods and watches them until they exit successfully. Once training succeeds, the Job is done and frees its resources.
apiVersion: batch/v1
kind: Job
metadata:
name: train-ranker
spec:
template:
spec:
restartPolicy: NeverrestartPolicy Must Not Be Always
A Job Pod must use restartPolicy Never or OnFailure. Always is for servers and Kubernetes rejects it for a Job that is meant to end.
Retries with backoffLimit
Training can fail on a flaky download. The backoffLimit sets how many times the Job retries a failing Pod before it gives up for good.
spec:
backoffLimit: 4
activeDeadlineSeconds: 3600Cap Runtime to Avoid Runaways
Set activeDeadlineSeconds so a hung training run cannot burn an expensive GPU node all weekend. The Job is killed once that limit passes. ⏱️
Request the GPU You Need
A training Job uses the same resources block as a Deployment. Add nvidia.com/gpu under limits so the run lands on a GPU node.
Parallelism for Sweeps
Set completions and parallelism to run many Pods, perfect for a hyperparameter sweep where each Pod trains one configuration at once.
Logs and Artifacts Outlive the Pod
A finished Job Pod is gone, so write your model and metrics to durable storage like S3 or a volume, never to the Pod filesystem.
CronJob for Scheduled Retraining
Wrap a Job in a CronJob to retrain on a schedule, like nightly. It creates a fresh Job each time the cron expression fires.
apiVersion: batch/v1
kind: CronJob
spec:
schedule: "0 2 * * *"Clean Up Finished Jobs
Completed Jobs linger by default and clutter the cluster. Set ttlSecondsAfterFinished so Kubernetes deletes them automatically after a grace period.
Watch Status with kubectl
Check a run with kubectl get jobs and read output with kubectl logs. The status shows succeeded or failed counts at a glance.
kubectl get jobs
kubectl logs job/train-rankerQuick Check
Why is a Job, not a Deployment, right for training?
Recap
Run finite training as a Job with restartPolicy Never, a backoffLimit, and a runtime cap, then schedule retraining with a CronJob. ✅
الأسئلة الشائعة
هل درس «تشغيل التدريب كـ Kubernetes Job» مجاني؟
نعم — نص درس «تشغيل التدريب كـ Kubernetes Job» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة MLOps Academy، انتقل إلى CoddyKit PRO. تتضمن دورة MLOps Academy 4 دروس في المجموع.
ماذا ستتعلم في «تشغيل التدريب كـ Kubernetes Job»؟
نفّذ التدريب الدفعي حتى اكتماله على العنقود. تتمرن على MLOps Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ MLOps Academy؟
لا تُشترط خبرة سابقة. MLOps Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.
كم من الوقت يستغرق درس «تشغيل التدريب كـ Kubernetes Job»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس MLOps Academy هذا؟
نعم. كل درس في MLOps Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- Pods وDeployments وServices للنماذج
- طلب CPU والذاكرة ووحدة GPU
- التهيئة باستخدام ConfigMaps وSecrets
- تشغيل التدريب كـ Kubernetes Job