تشغيل المهام باستخدام torchrun
أنشئ عمليات العاملين ونسّق بينها
تشغيل المهام باستخدام torchrun درس مجاني في Deep Learning Academy على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Deep Learning Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Deep Learning Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
Who Starts the Processes
DDP needs one process per GPU, but who spawns them? You do not launch them by hand. torchrun is the tool that does it for you.
Meet torchrun
torchrun is PyTorch's launcher. You give it your script and how many processes to start, and it handles the coordination.
torchrun --nproc_per_node=4 train.pyProcesses per Node
The nproc_per_node flag sets how many processes to spawn on this machine. Usually it equals the number of GPUs you have.
torchrun --nproc_per_node=8 train.pyEnvironment Variables Appear
torchrun injects key values as environment variables: RANK, LOCAL_RANK, and WORLD_SIZE. Your script reads them to know who it is.
import os
local_rank = int(os.environ["LOCAL_RANK"])No Manual Address Wiring
Older launchers made you pass ranks by hand. With torchrun, the rendezvous is automatic, so your script stays clean and portable.
Init Reads the Env
Because torchrun sets the env vars, your init_process_group call needs no arguments beyond the backend. It picks everything up automatically.
import torch.distributed as dist
dist.init_process_group(backend="nccl")Scale to Many Machines
To go multi-node, add nnodes and a node rank. Each machine runs the same command with its own node id.
torchrun --nnodes=2 --node_rank=0 --nproc_per_node=4 train.pyPoint to the Master
Across nodes, all processes must find a meeting point. The rdzv_endpoint gives the host and port they rendezvous on.
torchrun --rdzv_endpoint=host0:29500 train.pySurvive a Worker Crash
torchrun supports elastic training: set a min and max node count, and the job can recover if a worker drops out. 💪
torchrun --nnodes=1:4 --max_restarts=3 train.pyLog From One Rank
Every process prints, so logs get noisy. Gate prints behind a rank check so only rank 0 reports progress.
if int(os.environ["RANK"]) == 0:
print("epoch done")The Whole Recipe
The pattern is simple: write a normal DDP script, then launch it with torchrun. The launcher and DDP handle the rest together.
Quick Check
Recall what torchrun hands to your script.
Recap
You learned to launch jobs with torchrun: set processes per node, read the injected env vars, and scale to many nodes with elastic recovery.
تعلم Python مع معلم ذكاء اصطناعي — مجانًا
اكتب وقم بتشغيل أكوادك الفعلية في المتصفح، واحصل على مساعدة فورية من معلم ذكاء اصطناعي متاح 24/7، واستمر من حيث توقفت على الويب أو في التطبيق.
- الدورات
- 30
- الدروس
- 120
الأسئلة الشائعة
هل درس «تشغيل المهام باستخدام torchrun» مجاني؟
نعم — نص درس «تشغيل المهام باستخدام torchrun» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Deep Learning Academy، انتقل إلى CoddyKit PRO. تتضمن دورة Deep Learning Academy 4 دروس في المجموع.
ماذا ستتعلم في «تشغيل المهام باستخدام torchrun»؟
أنشئ عمليات العاملين ونسّق بينها تتمرن على Deep Learning Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ Deep Learning Academy؟
لا تُشترط خبرة سابقة. Deep Learning Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.
كم من الوقت يستغرق درس «تشغيل المهام باستخدام torchrun»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس Deep Learning Academy هذا؟
نعم. كل درس في Deep Learning Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- التوازي على مستوى البيانات مقابل النموذج
- أساسيات DistributedDataParallel
- مزامنة Batch Norm والحالة المجزأة
- تشغيل المهام باستخدام torchrun