Acesso ponto a ponto à memória
Cópias diretas de GPU para GPU pelo NVLink.
Acesso ponto a ponto à memória é uma aula grátis de CUDA Academy no CoddyKit. Esta é a aula 3 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de CUDA Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de CUDA Academy inclui 4 aulas no total.
Partes desta aula ainda não foram traduzidas e aparecem em inglês.
The Slow Detour
Moving data from one GPU to another usually bounces through the CPU's memory first. That round trip is slow and wastes the host's bandwidth.
GPUs Talking Directly
Modern GPUs can skip the CPU entirely. Peer-to-peer access lets one GPU read and write another's memory over a direct link.
The NVLink Highway
The fast link between cards is often NVLink, far quicker than the shared PCIe bus. P2P transfers ride this highway when it is available.
Check Before You Trust
Not every pair of GPUs can do P2P. Ask first with cudaDeviceCanAccessPeer, which reports whether one device may reach another.
int can;
cudaDeviceCanAccessPeer(&can, 0, 1);Turn On Access
Permission is off by default. While GPU 0 is current, call cudaDeviceEnablePeerAccess to let it reach into GPU 1's memory.
cudaSetDevice(0);
cudaDeviceEnablePeerAccess(1, 0);Access Is One-Way
Enabling peer access grants only the direction you ask for. For GPU 1 to read GPU 0, you must enable that direction too, from device 1.
Copying Peer to Peer
To move a buffer directly between cards, use cudaMemcpyPeer. You name the destination pointer and device plus the source pointer and device.
cudaMemcpyPeer(dst, 1, src, 0, bytes);Kernels Reading Across
With access enabled, a kernel on GPU 0 can dereference a pointer that lives on GPU 1. The hardware fetches it over the peer link transparently.
When P2P Is Not There
If two cards cannot peer, the runtime quietly falls back to staging through host memory. You still get a copy, just at the slower PCIe rate.
Async Peer Copies
Peer transfers can overlap other work. Pair cudaMemcpyPeerAsync with a stream so the copy runs while kernels keep computing.
cudaMemcpyPeerAsync(dst, 1, src, 0, bytes, stream);Turn It Off When Done
Peer access uses resources, so release it when finished. cudaDeviceDisablePeerAccess tears down the link you opened earlier.
cudaDeviceDisablePeerAccess(1);Quick Check
Recall the benefit of peer-to-peer access between two GPUs.
Recap
P2P lets GPUs share memory directly, ideally over NVLink. Check support, enable access per direction, then use cudaMemcpyPeer. Next: scaling with NCCL. ✨
Perguntas Frequentes
A aula “Acesso ponto a ponto à memória” é grátis?
Sim — o texto completo de “Acesso ponto a ponto à memória” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de CUDA Academy, atualize para CoddyKit PRO. O curso de CUDA Academy inclui 4 aulas no total.
O que vou aprender em “Acesso ponto a ponto à memória”?
Cópias diretas de GPU para GPU pelo NVLink. Você pratica CUDA Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.
Preciso ter experiência prévia para começar CUDA Academy?
Nenhuma experiência prévia é necessária. CUDA Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 3 de 4.
Quanto tempo leva a aula “Acesso ponto a ponto à memória”?
A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.
Posso escrever e executar código nesta aula de CUDA Academy?
Sim. Cada aula de CUDA Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.
Todas as aulas deste curso
- Enumerando e selecionando dispositivos
- Particionando o trabalho entre GPUs
- Acesso ponto a ponto à memória
- Várias GPUs com NCCL