0Pricing
CUDA Academy · درس

نواة Matmul الساذجة

خط أساس مفهرس ثنائي الأبعاد وحدوده.

نواة Matmul الساذجة درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 1 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

Matrix Multiply, GPU Style

Matrix multiplication is the heart of graphics and AI. Today you build a naive GPU version first, then learn why it leaves speed on the table.

The Math in One Line

Each output cell C[row][col] is a dot product: multiply a full row of A by a full column of B and sum the results. 🧮

C[row][col] = sum over k of A[row][k] * B[k][col]

One Thread per Output

The simplest plan gives each thread one output element of C. Thousands of cells get computed at the same time across the GPU.

A 2D Grid of Threads

Since C is a 2D grid, you launch threads in two dimensions. The x index maps to a column and the y index maps to a row.

dim3 threads(16, 16);
dim3 blocks((N+15)/16, (N+15)/16);

Finding This Thread's Cell

Inside the kernel, each thread computes its own row and col from its block and thread indices, just like 1D indexing but on both axes.

int row = blockIdx.y*blockDim.y + threadIdx.y;
int col = blockIdx.x*blockDim.x + threadIdx.x;

The Bounds Check

Grids round up, so some threads fall outside the matrix. Guard with if (row < N && col < N) before you touch memory.

if (row < N && col < N) {
  // safe to compute
}

The Inner Loop

Each thread runs a loop over k, accumulating products into a local sum. That local variable lives in a fast register.

float sum = 0.0f;
for (int k = 0; k < N; ++k)
  sum += A[row*N+k] * B[k*N+col];

Writing the Result

After the loop finishes, the thread stores its accumulated sum into C exactly once. One thread, one clean write.

C[row*N + col] = sum;

Row-Major Flattening

The matrix is a flat 1D array, so you index it as row*N + col. Getting this layout right is half the battle in matmul.

Why It Works, But Slowly

This kernel is correct and easy to read, but every thread reads its row and column straight from global memory, the slowest space.

The Hidden Cost

Neighboring threads re-read the same A rows and B columns over and over. That wasted memory traffic is exactly what tiling will fix next.

Quick Check

Think about how the naive kernel maps work to threads.

Recap

You mapped one thread to one output cell, looped over k from global memory, and saw the redundant reads. Next you cut that traffic with tiling. 🚀

الأسئلة الشائعة

هل درس «نواة Matmul الساذجة» مجاني؟

نعم — نص درس «نواة Matmul الساذجة» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.

ماذا ستتعلم في «نواة Matmul الساذجة»؟

خط أساس مفهرس ثنائي الأبعاد وحدوده. تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟

لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 1 من أصل 4.

كم من الوقت يستغرق درس «نواة Matmul الساذجة»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟

نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. نواة Matmul الساذجة
  2. تقسيم الجداء الداخلي إلى بلاطات
  3. التكرار عبر مراحل البلاطات
  4. قياس التسارع
← العودة إلى CUDA Academy