0Pricing
CUDA Academy · Aula

Leituras coalescidas versus entrelaçadas

Veja como o mapeamento entre threads e endereços importa.

Leituras coalescidas versus entrelaçadas é uma aula grátis de CUDA Academy no CoddyKit. Esta é a aula 2 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de CUDA Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de CUDA Academy inclui 4 aulas no total.

Partes desta aula ainda não foram traduzidas e aparecem em inglês.

It Is All About Mapping

Coalescing depends on which address each thread touches. The map from thread to address decides whether one transaction serves the whole warp.

The Coalesced Pattern

When neighboring threads read neighboring addresses, the warp covers one contiguous run. That tidy layout is a coalesced access.

int i = blockIdx.x * blockDim.x + threadIdx.x;
float v = data[i];

Why It Wins

Thread 0 hits index 0, thread 1 hits index 1, and so on. All 32 addresses fall in one aligned line, so the GPU needs a single transaction.

Enter the Stride

A stride is a fixed gap between the addresses successive threads touch. The moment that gap grows past one element, coalescing starts to break.

float v = data[i * stride];

Strided Reads Spread Out

With a stride of 2, thread 0 reads 0, thread 1 reads 2, thread 2 reads 4. The warp now spans twice the memory, so it needs more transactions.

The Cost Scales Up

Double the stride and you roughly double the lines touched. Big strides can push a warp toward many transactions while using only a sliver of each line.

A Common Trap

Giving each thread a whole column of a row-major matrix creates a huge stride. It looks neat in code but quietly scatters every warp.

float v = matrix[threadIdx.x * width + row];

Transpose the Mapping

Often the fix is just swapping which index varies fastest with the thread. Let consecutive threads walk consecutive memory and the reads coalesce again.

float v = matrix[row * width + threadIdx.x];

Offsets Hurt Too

Even a coalesced pattern suffers if it starts mid-line. A misaligned offset shifts the warp across a boundary, splitting one read into two.

Think in Warps

To judge a pattern, do not picture one thread. Picture all 32 lanes at once and ask how many lines their addresses cover together. 🔍

The Rule of Thumb

Make the fastest-changing index follow threadIdx.x. That single habit keeps most of your global reads coalesced for free.

Quick Check

Pick the access pattern that coalesces best.

Recap

You saw that coalesced reads keep neighbors together while a stride scatters them across lines. Keep threadIdx.x driving the fastest index. 🎉

Perguntas Frequentes

A aula “Leituras coalescidas versus entrelaçadas” é grátis?

Sim — o texto completo de “Leituras coalescidas versus entrelaçadas” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de CUDA Academy, atualize para CoddyKit PRO. O curso de CUDA Academy inclui 4 aulas no total.

O que vou aprender em “Leituras coalescidas versus entrelaçadas”?

Veja como o mapeamento entre threads e endereços importa. Você pratica CUDA Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.

Preciso ter experiência prévia para começar CUDA Academy?

Nenhuma experiência prévia é necessária. CUDA Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 2 de 4.

Quanto tempo leva a aula “Leituras coalescidas versus entrelaçadas”?

A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.

Posso escrever e executar código nesta aula de CUDA Academy?

Sim. Cada aula de CUDA Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.

Todas as aulas deste curso

  1. O que é uma transação de memória
  2. Leituras coalescidas versus entrelaçadas
  3. Estrutura de vetores versus vetor de estruturas
  4. Meça a largura de banda efetiva
← Voltar para CUDA Academy