Memória fixada com cudaMallocHost
Buffers com páginas bloqueadas para DMA rápido.
Memória fixada com cudaMallocHost é uma aula grátis de CUDA Academy no CoddyKit. Esta é a aula 2 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de CUDA Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de CUDA Academy inclui 4 aulas no total.
Partes desta aula ainda não foram traduzidas e aparecem em inglês.
Meet Pinned Memory
Pinned memory is host memory the operating system promises never to move or swap out. Because its physical address is fixed, the GPU can read it directly.
Also Called Page-Locked
You will see the term page-locked used as a synonym for pinned. Both mean the same thing: pages locked in place so no swapping can ever happen.
Allocate It Right
Ask CUDA for pinned memory with cudaMallocHost. It hands back an ordinary host pointer you can read and write just like any C++ array.
float* h_data;
cudaMallocHost(&h_data, N * sizeof(float));Use It Like Normal
The returned pointer behaves like any host buffer, so your CPU code stays unchanged. You just fill it and read it exactly as before.
for (int i = 0; i < N; ++i) h_data[i] = i * 1.0f;No More Staging
Since the pages cannot move, the driver skips its hidden staging copy. The GPU pulls straight from your buffer, so the doubled copy disappears.
Faster Transfers
The payoff is real bandwidth. Pinned transfers often run near the link's peak speed, sometimes roughly twice as fast as the pageable version.
Free It Correctly
Pinned memory needs its own matching release. Use cudaFreeHost, never the plain free, or you risk a crash and a leak.
cudaFreeHost(h_data);It Is Not Free Real Estate
Pinned pages cannot be swapped, so they hold real RAM hostage. Pin too much and you starve the rest of the system of usable memory.
Allocation Is Slower
Locking pages takes work, so cudaMallocHost is slower to allocate than malloc. Reuse one pinned buffer across many transfers instead of reallocating.
The Real Superpower
Beyond speed, pinned memory is the key that unlocks truly asynchronous copies. Without it, cudaMemcpyAsync cannot overlap with anything.
When to Reach for It
Pin buffers that you transfer often or that feed streams. For a one-off copy of a tiny array, plain malloc is perfectly fine. 🙂
Quick Check
You allocated a buffer with cudaMallocHost. How should you release it?
Recap
cudaMallocHost pins host pages so the GPU reads them directly, giving faster copies and enabling async overlap. Always release with cudaFreeHost. 🔒
Perguntas Frequentes
A aula “Memória fixada com cudaMallocHost” é grátis?
Sim — o texto completo de “Memória fixada com cudaMallocHost” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de CUDA Academy, atualize para CoddyKit PRO. O curso de CUDA Academy inclui 4 aulas no total.
O que vou aprender em “Memória fixada com cudaMallocHost”?
Buffers com páginas bloqueadas para DMA rápido. Você pratica CUDA Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.
Preciso ter experiência prévia para começar CUDA Academy?
Nenhuma experiência prévia é necessária. CUDA Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 2 de 4.
Quanto tempo leva a aula “Memória fixada com cudaMallocHost”?
A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.
Posso escrever e executar código nesta aula de CUDA Academy?
Sim. Cada aula de CUDA Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.
Todas as aulas deste curso
- Por que a memória paginável é lenta
- Memória fixada com cudaMallocHost
- cudaMemcpyAsync em um fluxo
- O pipeline de buffers duplos