0Pricing
CUDA Academy · Ders

Saf Matris Çarpımı Çekirdeği

İki boyutlu dizin kullanan temel uygulama ve sınırları.

Saf Matris Çarpımı Çekirdeği, CoddyKit'te ücretsiz bir CUDA Academy dersidir. Bu, 4 dersinin 1. dersidir. Aşağıdan dersin tamamını ücretsiz okuyabilir, sonra tarayıcıda yerleşik kod editörü ve 7/24 yapay zeka koçu ile uygulamalı olarak pratik yapabilirsin. Bu, CUDA Academy öğrenme yolunun bir parçasıdır ve ilerlemeniz web ve CoddyKit uygulaması arasında senkronize olur. CUDA Academy kursu toplamda 4 dersten oluşur.

Bu dersin bazı bölümleri henüz çevrilmemiş olup İngilizce olarak gösterilmektedir.

Matrix Multiply, GPU Style

Matrix multiplication is the heart of graphics and AI. Today you build a naive GPU version first, then learn why it leaves speed on the table.

The Math in One Line

Each output cell C[row][col] is a dot product: multiply a full row of A by a full column of B and sum the results. 🧮

C[row][col] = sum over k of A[row][k] * B[k][col]

One Thread per Output

The simplest plan gives each thread one output element of C. Thousands of cells get computed at the same time across the GPU.

A 2D Grid of Threads

Since C is a 2D grid, you launch threads in two dimensions. The x index maps to a column and the y index maps to a row.

dim3 threads(16, 16);
dim3 blocks((N+15)/16, (N+15)/16);

Finding This Thread's Cell

Inside the kernel, each thread computes its own row and col from its block and thread indices, just like 1D indexing but on both axes.

int row = blockIdx.y*blockDim.y + threadIdx.y;
int col = blockIdx.x*blockDim.x + threadIdx.x;

The Bounds Check

Grids round up, so some threads fall outside the matrix. Guard with if (row < N && col < N) before you touch memory.

if (row < N && col < N) {
  // safe to compute
}

The Inner Loop

Each thread runs a loop over k, accumulating products into a local sum. That local variable lives in a fast register.

float sum = 0.0f;
for (int k = 0; k < N; ++k)
  sum += A[row*N+k] * B[k*N+col];

Writing the Result

After the loop finishes, the thread stores its accumulated sum into C exactly once. One thread, one clean write.

C[row*N + col] = sum;

Row-Major Flattening

The matrix is a flat 1D array, so you index it as row*N + col. Getting this layout right is half the battle in matmul.

Why It Works, But Slowly

This kernel is correct and easy to read, but every thread reads its row and column straight from global memory, the slowest space.

The Hidden Cost

Neighboring threads re-read the same A rows and B columns over and over. That wasted memory traffic is exactly what tiling will fix next.

Quick Check

Think about how the naive kernel maps work to threads.

Recap

You mapped one thread to one output cell, looped over k from global memory, and saw the redundant reads. Next you cut that traffic with tiling. 🚀

Sıkça Sorulan Sorular

“Saf Matris Çarpımı Çekirdeği” dersi ücretsiz mi?

Evet — “Saf Matris Çarpımı Çekirdeği” dersin tüm metni burada web'de ücretsiz olarak okunabilir. Etkileşimli olarak pratik yapmak (yerleşik kod editörü ve 7/24 yapay zeka koçu) ve CUDA Academy kursunun geri kalanını açmak için CoddyKit PRO'ya yükselt. CUDA Academy kursu toplamda 4 dersten oluşur.

“Saf Matris Çarpımı Çekirdeği” dersinde ne öğreneceğim?

İki boyutlu dizin kullanan temel uygulama ve sınırları. CUDA Academy ile uygulamalı kodu tarayıcıda doğrudan çalıştırarak pratik yaparsın ve 7/24 yapay zeka koçu dersi çalışırken sorularını yanıtlar.

CUDA Academy öğrenmeye başlamak için deneyim gerekli mi?

Önceden deneyim gerekmez. CoddyKit'te CUDA Academy, başlangıçtan ileri seviyeye kadar yapılandırıldığı için buradan başlayabilir veya başından başlayıp kendi hızında ilerleme yapabilirsin. Bu, 4 dersinin 1. dersidir.

“Saf Matris Çarpımı Çekirdeği” dersi ne kadar sürer?

Çoğu CoddyKit dersi yaklaşık 5–10 dakika sürer. Her biri kısa ve etkileşimli olduğu için sabit ilerleme yaparsın ve web ile uygulama arasında tam olarak bıraktığın yerden devam edebilirsin.

Bu CUDA Academy dersinde kod yazıp çalıştırabilir miyim?

Evet. Her CUDA Academy dersi yerleşik bir kod editörü içerir, bu sayede tarayıcıda gerçek kod yazıp çalıştırabilir ve anlık yapay zeka geri bildirimi alırsın — yerel kurulum gerekli değildir.

Bu kursun tüm dersleri

  1. Saf Matris Çarpımı Çekirdeği
  2. İç Çarpımı Döşeyin
  3. Döşeme Aşamaları Üzerinde Döngü
  4. Hızlanmayı Ölçün
← CUDA Academy Sayfasına Dön