0Pricing
Mojo Academy · درس

بناء Matmul خطوة بخطوة

من الحلقات الساذجة إلى نواة حقيقية

بناء Matmul خطوة بخطوة درس مجاني في Mojo Academy على CoddyKit. هذا هو الدرس 2 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Mojo Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Mojo Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

What Matmul Computes

Matrix multiply, or matmul, combines an M by K matrix with a K by N matrix to produce an M by N result. It is the heart of AI math.

The Dot Product Rule

Each output cell is a dot product: walk across one row of A and down one column of B, multiplying and summing as you go.

Three Nested Loops

The naive version uses three loops over i, j, and k. The outer two pick the output cell, the inner one sums the products.

for i in range(M):
    for j in range(N):
        for k in range(K):
            C[i, j] += A[i, k] * B[k, j]

Start Each Cell at Zero

Before accumulating, set each result cell to zero. Otherwise old or garbage values pollute your sum.

C[i, j] = 0.0

Accumulate the Sum

The inner k loop keeps a running accumulator. Summing into a local variable is often faster than touching C every step.

var acc: Float32 = 0.0
for k in range(K):
    acc += A[i, k] * B[k, j]
C[i, j] = acc

Use fn for Speed

Write the kernel with fn and typed arguments. Strict types let Mojo compile tight machine code with no dynamic overhead.

fn matmul(A: Matrix, B: Matrix, C: Matrix):
    pass

Why Naive Is Slow

The basic triple loop does the right math but reads B by column, jumping through memory. Poor locality wastes cache and time.

Loop Order Matters

Reordering to i, k, j keeps the inner loop walking memory in straight lines. Better access patterns can speed matmul a lot.

for i in range(M):
    for k in range(K):
        for j in range(N):
            C[i, j] += A[i, k] * B[k, j]

The Inner Loop Is the Target

Nearly all the time lives in the innermost loop. That hot inner loop is exactly where vectorizing and tuning pay off.

Correctness First

Get the simple version right and save its output. It becomes the reference you compare every faster kernel against.

A Path to a Real Kernel

From here you add SIMD, tiling, and parallelism. Each step keeps the same result but raises throughput toward peak hardware speed.

Quick Check

Why is the textbook triple-loop matmul often slow in practice?

Recap

Matmul sums a dot product per output cell with three loops; start cells at zero, accumulate locally, and mind loop order for cache. 🔢

الأسئلة الشائعة

هل درس «بناء Matmul خطوة بخطوة» مجاني؟

نعم — نص درس «بناء Matmul خطوة بخطوة» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Mojo Academy، انتقل إلى CoddyKit PRO. تتضمن دورة Mojo Academy 4 دروس في المجموع.

ماذا ستتعلم في «بناء Matmul خطوة بخطوة»؟

من الحلقات الساذجة إلى نواة حقيقية تتمرن على Mojo Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ Mojo Academy؟

لا تُشترط خبرة سابقة. Mojo Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 2 من أصل 4.

كم من الوقت يستغرق درس «بناء Matmul خطوة بخطوة»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس Mojo Academy هذا؟

نعم. كل درس في Mojo Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. نمذجة موتر في Mojo
  2. بناء Matmul خطوة بخطوة
  3. تحسين الجداء الداخلي
  4. التحقق من الصحة العددية
← العودة إلى Mojo Academy