0Pricing
CUDA Academy · Aula

Divida o produto interno em blocos

Carregue subblocos de A e B a cada fase.

Divida o produto interno em blocos é uma aula grátis de CUDA Academy no CoddyKit. Esta é a aula 2 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de CUDA Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de CUDA Academy inclui 4 aulas no total.

Partes desta aula ainda não foram traduzidas e aparecem em inglês.

The Tiling Idea

Tiling breaks the matrices into small square tiles that fit in fast on-chip memory. Threads cooperate to load a tile once and reuse it many times.

Why Shared Memory

A tile lives in __shared__ memory, visible to every thread in the block. Reading it is far faster than hitting global memory again and again. ⚡

__shared__ float As[TILE][TILE];
__shared__ float Bs[TILE][TILE];

One Block, One Output Tile

Each block is responsible for one TILE-by-TILE patch of the output C. Its threads team up to compute that whole patch together.

Tile Size Matches Block

You pick TILE equal to the block's width, often 16 or 32. That way each thread loads exactly one element of each tile.

#define TILE 16
dim3 threads(TILE, TILE);

Mapping Thread to Tile Slot

Inside the tile, a thread's slot is just its threadIdx. Its global row and col still come from the block and thread indices.

int ty = threadIdx.y, tx = threadIdx.x;
int row = blockIdx.y*TILE + ty;
int col = blockIdx.x*TILE + tx;

Loading a Tile of A

Each thread copies one element of A's current tile into shared memory. Together the block stages a full TILE-by-TILE block of A.

As[ty][tx] = A[row*N + (phase*TILE + tx)];

Loading a Tile of B

At the same time, each thread loads one element of B's tile. Now both tiles sit on chip, ready for fast repeated reads.

Bs[ty][tx] = B[(phase*TILE + ty)*N + col];

Sync Before You Compute

Call __syncthreads() so every thread finishes loading before anyone reads the tile. Skipping this gives garbage results.

__syncthreads();

Compute on the Tile

Now each thread does a short loop over the tile, reading only shared memory. These reads are dramatically cheaper than global ones.

for (int k = 0; k < TILE; ++k)
  sum += As[ty][k] * Bs[k][tx];

The Reuse Payoff

Every loaded value gets used by TILE threads instead of one. That data reuse is the whole reason tiled matmul flies.

One Tile Is Not Enough

A single tile only covers part of the dot product. You repeat the load-sync-compute steps across many phases, which the next lesson handles.

Quick Check

Recall the order of steps when working with a shared tile.

Recap

You staged tiles of A and B into shared memory, synced, then computed with cheap on-chip reads. Reuse is the win. Phases come next. 🧱

Perguntas Frequentes

A aula “Divida o produto interno em blocos” é grátis?

Sim — o texto completo de “Divida o produto interno em blocos” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de CUDA Academy, atualize para CoddyKit PRO. O curso de CUDA Academy inclui 4 aulas no total.

O que vou aprender em “Divida o produto interno em blocos”?

Carregue subblocos de A e B a cada fase. Você pratica CUDA Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.

Preciso ter experiência prévia para começar CUDA Academy?

Nenhuma experiência prévia é necessária. CUDA Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 2 de 4.

Quanto tempo leva a aula “Divida o produto interno em blocos”?

A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.

Posso escrever e executar código nesta aula de CUDA Academy?

Sim. Cada aula de CUDA Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.

Todas as aulas deste curso

  1. O núcleo ingênuo de multiplicação de matrizes
  2. Divida o produto interno em blocos
  3. Percorra as fases dos blocos
  4. Meça o ganho de velocidade
← Voltar para CUDA Academy