DistributedDataParallel Basics
The standard multi-GPU training path.
DistributedDataParallel Basics is a free Deep Learning Academy lesson on CoddyKit — lesson 2 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Deep Learning Academy learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
Meet DDP
DistributedDataParallel, or DDP, is PyTorch's go-to tool for multi-GPU training. It runs one process per GPU and keeps every model copy in sync.
One Process per GPU
Unlike the older DataParallel, DDP spawns a separate process for each GPU. This avoids Python's GIL and scales far more cleanly.
Rank and World Size
Each process gets a rank (its id) and shares the world size (total processes). Rank 0 is usually the one that logs and saves.
import torch.distributed as dist
rank = dist.get_rank()
world = dist.get_world_size()Init the Process Group
Before any communication you call init_process_group. The nccl backend is the fast choice for GPUs.
import torch.distributed as dist
dist.init_process_group(backend="nccl")Pin Each Process to a GPU
Use the local rank to set the device so every process owns exactly one GPU. This keeps work from piling onto a single card.
import torch
torch.cuda.set_device(local_rank)
model = model.to(local_rank)Wrap Your Model
The magic is one line: wrap your model in DDP. From then on, gradients sync automatically during the backward pass.
from torch.nn.parallel import DistributedDataParallel as DDP
model = DDP(model, device_ids=[local_rank])Gradients Sync Themselves
During backward(), DDP performs an all-reduce to average gradients across GPUs. You write normal training code and it just stays in sync. ✨
Use a DistributedSampler
So each GPU sees different data, give your DataLoader a DistributedSampler. It hands every process a non-overlapping slice of the dataset.
from torch.utils.data.distributed import DistributedSampler
sampler = DistributedSampler(dataset)Reshuffle Every Epoch
Call sampler.set_epoch(epoch) at the top of each epoch. Without it, every GPU reshuffles the same way and you lose real shuffling.
for epoch in range(epochs):
sampler.set_epoch(epoch)
train_one_epoch()Save Only on Rank 0
All copies are identical, so checkpoint from rank 0 only. Saving from every process just writes the same file many times.
if rank == 0:
torch.save(model.module.state_dict(), "ckpt.pt")Clean Up at the End
When training finishes, call destroy_process_group to release the group cleanly and avoid hanging processes.
import torch.distributed as dist
dist.destroy_process_group()Quick Check
Think about how DDP keeps copies in sync.
Recap
You set up DDP: init the process group, wrap the model, feed it a DistributedSampler, and save from rank 0. Gradients sync for free.
Frequently asked questions
Is the “DistributedDataParallel Basics” lesson free?
Yes — the full text of “DistributedDataParallel Basics” is free to read here on the web, and the Deep Learning Academy course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Deep Learning Academy course, upgrade to CoddyKit PRO.
What will I learn in “DistributedDataParallel Basics”?
The standard multi-GPU training path. You practise Deep Learning Academy with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start Deep Learning Academy?
No prior experience is required. Deep Learning Academy on CoddyKit is structured for beginners through advanced learners; this is — lesson 2 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “DistributedDataParallel Basics” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this Deep Learning Academy lesson?
Yes. Every Deep Learning Academy lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.
All lessons in this course
- Data vs Model Parallelism
- DistributedDataParallel Basics
- Sync Batch Norm & Sharded State
- Launch Jobs with torchrun