0Pricing
Deep Learning Academy · Aula

Fundamentos de DistributedDataParallel

O caminho padrão para treinamento com várias GPUs

Fundamentos de DistributedDataParallel é uma aula grátis de Deep Learning Academy no CoddyKit. Esta é a aula 2 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de Deep Learning Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de Deep Learning Academy inclui 4 aulas no total.

Partes desta aula ainda não foram traduzidas e aparecem em inglês.

Meet DDP

DistributedDataParallel, or DDP, is PyTorch's go-to tool for multi-GPU training. It runs one process per GPU and keeps every model copy in sync.

One Process per GPU

Unlike the older DataParallel, DDP spawns a separate process for each GPU. This avoids Python's GIL and scales far more cleanly.

Rank and World Size

Each process gets a rank (its id) and shares the world size (total processes). Rank 0 is usually the one that logs and saves.

import torch.distributed as dist
rank = dist.get_rank()
world = dist.get_world_size()

Init the Process Group

Before any communication you call init_process_group. The nccl backend is the fast choice for GPUs.

import torch.distributed as dist
dist.init_process_group(backend="nccl")

Pin Each Process to a GPU

Use the local rank to set the device so every process owns exactly one GPU. This keeps work from piling onto a single card.

import torch
torch.cuda.set_device(local_rank)
model = model.to(local_rank)

Wrap Your Model

The magic is one line: wrap your model in DDP. From then on, gradients sync automatically during the backward pass.

from torch.nn.parallel import DistributedDataParallel as DDP
model = DDP(model, device_ids=[local_rank])

Gradients Sync Themselves

During backward(), DDP performs an all-reduce to average gradients across GPUs. You write normal training code and it just stays in sync. ✨

Use a DistributedSampler

So each GPU sees different data, give your DataLoader a DistributedSampler. It hands every process a non-overlapping slice of the dataset.

from torch.utils.data.distributed import DistributedSampler
sampler = DistributedSampler(dataset)

Reshuffle Every Epoch

Call sampler.set_epoch(epoch) at the top of each epoch. Without it, every GPU reshuffles the same way and you lose real shuffling.

for epoch in range(epochs):
    sampler.set_epoch(epoch)
    train_one_epoch()

Save Only on Rank 0

All copies are identical, so checkpoint from rank 0 only. Saving from every process just writes the same file many times.

if rank == 0:
    torch.save(model.module.state_dict(), "ckpt.pt")

Clean Up at the End

When training finishes, call destroy_process_group to release the group cleanly and avoid hanging processes.

import torch.distributed as dist
dist.destroy_process_group()

Quick Check

Think about how DDP keeps copies in sync.

Recap

You set up DDP: init the process group, wrap the model, feed it a DistributedSampler, and save from rank 0. Gradients sync for free.

Perguntas Frequentes

A aula “Fundamentos de DistributedDataParallel” é grátis?

Sim — o texto completo de “Fundamentos de DistributedDataParallel” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de Deep Learning Academy, atualize para CoddyKit PRO. O curso de Deep Learning Academy inclui 4 aulas no total.

O que vou aprender em “Fundamentos de DistributedDataParallel”?

O caminho padrão para treinamento com várias GPUs Você pratica Deep Learning Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.

Preciso ter experiência prévia para começar Deep Learning Academy?

Nenhuma experiência prévia é necessária. Deep Learning Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 2 de 4.

Quanto tempo leva a aula “Fundamentos de DistributedDataParallel”?

A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.

Posso escrever e executar código nesta aula de Deep Learning Academy?

Sim. Cada aula de Deep Learning Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.

Todas as aulas deste curso

  1. Paralelismo de Dados vs de Modelo
  2. Fundamentos de DistributedDataParallel
  3. Normalização de Lote Sincronizada e Estado Fragmentado
  4. Inicie Tarefas com torchrun
← Voltar para Deep Learning Academy