Deep Learning Academy · درس

تشغيل المهام باستخدام torchrun

أنشئ عمليات العاملين ونسّق بينها

الدرس 4 من 413 خطوة

تشغيل المهام باستخدام torchrun درس مجاني في Deep Learning Academy على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Deep Learning Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Deep Learning Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

Who Starts the Processes

DDP needs one process per GPU, but who spawns them? You do not launch them by hand. torchrun is the tool that does it for you.

Meet torchrun

torchrun is PyTorch's launcher. You give it your script and how many processes to start, and it handles the coordination.

torchrun --nproc_per_node=4 train.py

Processes per Node

The nproc_per_node flag sets how many processes to spawn on this machine. Usually it equals the number of GPUs you have.

torchrun --nproc_per_node=8 train.py

Environment Variables Appear

torchrun injects key values as environment variables: RANK, LOCAL_RANK, and WORLD_SIZE. Your script reads them to know who it is.

import os
local_rank = int(os.environ["LOCAL_RANK"])

No Manual Address Wiring

Older launchers made you pass ranks by hand. With torchrun, the rendezvous is automatic, so your script stays clean and portable.

Init Reads the Env

Because torchrun sets the env vars, your init_process_group call needs no arguments beyond the backend. It picks everything up automatically.

import torch.distributed as dist
dist.init_process_group(backend="nccl")

Scale to Many Machines

To go multi-node, add nnodes and a node rank. Each machine runs the same command with its own node id.

torchrun --nnodes=2 --node_rank=0 --nproc_per_node=4 train.py

Point to the Master

Across nodes, all processes must find a meeting point. The rdzv_endpoint gives the host and port they rendezvous on.

torchrun --rdzv_endpoint=host0:29500 train.py

Survive a Worker Crash

torchrun supports elastic training: set a min and max node count, and the job can recover if a worker drops out. 💪

torchrun --nnodes=1:4 --max_restarts=3 train.py

Log From One Rank

Every process prints, so logs get noisy. Gate prints behind a rank check so only rank 0 reports progress.

if int(os.environ["RANK"]) == 0:
    print("epoch done")

The Whole Recipe

The pattern is simple: write a normal DDP script, then launch it with torchrun. The launcher and DDP handle the rest together.

Quick Check

Recall what torchrun hands to your script.

Recap

You learned to launch jobs with torchrun: set processes per node, read the injected env vars, and scale to many nodes with elastic recovery.

البدء مجانًا

تعلم Python مع معلم ذكاء اصطناعي — مجانًا

اكتب وقم بتشغيل أكوادك الفعلية في المتصفح، واحصل على مساعدة فورية من معلم ذكاء اصطناعي متاح 24/7، واستمر من حيث توقفت على الويب أو في التطبيق.

الدورات
30
الدروس
120

الأسئلة الشائعة

هل درس «تشغيل المهام باستخدام torchrun» مجاني؟

نعم — نص درس «تشغيل المهام باستخدام torchrun» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Deep Learning Academy، انتقل إلى CoddyKit PRO. تتضمن دورة Deep Learning Academy 4 دروس في المجموع.

ماذا ستتعلم في «تشغيل المهام باستخدام torchrun»؟

أنشئ عمليات العاملين ونسّق بينها تتمرن على Deep Learning Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ Deep Learning Academy؟

لا تُشترط خبرة سابقة. Deep Learning Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.

كم من الوقت يستغرق درس «تشغيل المهام باستخدام torchrun»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس Deep Learning Academy هذا؟

نعم. كل درس في Deep Learning Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. التوازي على مستوى البيانات مقابل النموذج
  2. أساسيات DistributedDataParallel
  3. مزامنة Batch Norm والحالة المجزأة
  4. تشغيل المهام باستخدام torchrun
← العودة إلى Deep Learning Academy