Um modelo mental da hierarquia
Associe os dados ao espaço adequado.
Um modelo mental da hierarquia é uma aula grátis de CUDA Academy no CoddyKit. Esta é a aula 4 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de CUDA Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de CUDA Academy inclui 4 aulas no total.
Partes desta aula ainda não foram traduzidas e aparecem em inglês.
A Pyramid of Tradeoffs
Think of GPU memory as a pyramid. The top is tiny and lightning fast, while the base is huge but slow. Your job is to use each layer well.
Top: Registers
At the peak sit registers: per-thread, fastest, and very limited. Keep your hottest working values here whenever you possibly can.
Next: Shared Memory
Just below comes shared memory, an on-chip scratchpad visible to all threads in a block. It is your tool for fast intra-block teamwork.
Side Path: Constant Cache
Alongside sits the constant cache, ideal for small read-only values that a warp reads uniformly and broadcasts in one fetch.
Base: Global Memory
The wide base is global memory: gigabytes, visible to everyone, but high-latency. It holds your big inputs and outputs.
The Trap: Local Memory
Watch out for local memory. It sounds close by but lives in slow DRAM, used only when registers spill. Avoid relying on it.
Scope Decides the Space
Pick by who needs the data. One thread alone wants registers, a block working together wants shared memory, everyone wants global.
Lifetime Decides Too
Registers and shared memory vanish when a kernel ends, but global memory persists across launches. Match storage to how long data must live.
The Golden Rule
The winning pattern is load once from global, compute in fast on-chip memory, then write once back. This minimizes slow global memory traffic.
Capacity Costs Occupancy
Spending lots of registers or shared memory per block lets fewer blocks run at once. Occupancy is the balance you constantly tune.
Putting It Together
Great kernels deliberately route each piece of data to the right layer. That single habit is what separates slow code from fast CUDA code. 💪
Quick Check
Two threads in the SAME block need to share intermediate results. Which space fits best?
Recap: Match Data to Layer
You built a mental map: registers, shared, constant, and global trade speed for size and scope. Routing data to the right layer is the whole game. 🧠
Perguntas Frequentes
A aula “Um modelo mental da hierarquia” é grátis?
Sim — o texto completo de “Um modelo mental da hierarquia” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de CUDA Academy, atualize para CoddyKit PRO. O curso de CUDA Academy inclui 4 aulas no total.
O que vou aprender em “Um modelo mental da hierarquia”?
Associe os dados ao espaço adequado. Você pratica CUDA Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.
Preciso ter experiência prévia para começar CUDA Academy?
Nenhuma experiência prévia é necessária. CUDA Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 4 de 4.
Quanto tempo leva a aula “Um modelo mental da hierarquia”?
A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.
Posso escrever e executar código nesta aula de CUDA Academy?
Sim. Cada aula de CUDA Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.
Todas as aulas deste curso
- Registradores e memória local
- Compromissos da memória global
- Memória constante e seu cache
- Um modelo mental da hierarquia