0Pricing
CUDA Academy · Aula

Laços com passo da grade

Trate vetores maiores que a grade.

Laços com passo da grade é uma aula grátis de CUDA Academy no CoddyKit. Esta é a aula 4 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de CUDA Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de CUDA Academy inclui 4 aulas no total.

Partes desta aula ainda não foram traduzidas e aparecem em inglês.

When the Array Is Huge

Sometimes the data is far bigger than the threads you launch. One thread per element no longer fits, so you need each thread to handle several elements.

The Grid Has a Size

The total number of threads is the grid size: blocks times threads per block. This stride is how far apart each thread's elements sit.

int stride = blockDim.x * gridDim.x;

Start, Then Stride

Each thread begins at its usual global index, then jumps forward by the grid size again and again until it runs off the array.

The Grid-Stride Loop

This loop is the whole pattern: start at i, step by stride, stop at n. Any array size is covered no matter how many threads you launch.

for (int i = blockIdx.x * blockDim.x + threadIdx.x;
     i < n;
     i += stride) {
    out[i] = a[i] + b[i];
}

The Built-in Guard

Notice the loop condition i < n is itself the bounds check. Threads that start past the end simply never enter the loop.

How the Work Splits

With a stride of 4 threads, thread 0 does elements 0, 4, 8, while thread 1 does 1, 5, 9. The work is interleaved, not chunked.

Interleaving Helps Coalescing

Because neighboring threads still touch neighboring addresses each step, the access stays coalesced and memory bandwidth stays high.

Decouple Threads From Data

Now your launch size no longer depends on n. You can pick a thread count that fits the GPU and let the loop absorb any workload.

Tune for the Hardware

A common choice is enough blocks to fill every multiprocessor, then let each thread loop. This keeps the GPU busy without overlaunching.

It Also Works When Tiny

If n is smaller than the grid, each thread runs the loop body at most once. The pattern degrades gracefully to the simple case.

A Robust Default

Many CUDA pros write every elementwise kernel as a grid-stride loop. It is flexible, safe, and rarely the wrong choice. 🚀

Quick Check

Identify the stride.

Recap

You learned the grid-stride loop: start at the global index and step by blockDim.x * gridDim.x until i reaches n. One kernel now handles any array size. 🎉

Perguntas Frequentes

A aula “Laços com passo da grade” é grátis?

Sim — o texto completo de “Laços com passo da grade” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de CUDA Academy, atualize para CoddyKit PRO. O curso de CUDA Academy inclui 4 aulas no total.

O que vou aprender em “Laços com passo da grade”?

Trate vetores maiores que a grade. Você pratica CUDA Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.

Preciso ter experiência prévia para começar CUDA Academy?

Nenhuma experiência prévia é necessária. CUDA Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 4 de 4.

Quanto tempo leva a aula “Laços com passo da grade”?

A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.

Posso escrever e executar código nesta aula de CUDA Academy?

Sim. Cada aula de CUDA Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.

Todas as aulas deste curso

  1. A fórmula clássica de índice
  2. Proteção contra índices fora do intervalo
  3. Arredonde o número de blocos para cima
  4. Laços com passo da grade
← Voltar para CUDA Academy