0Pricing
Mojo Academy · درس

دمج SIMD مع الحلقات

وجّه جوهر النواة باستخدام المتجهات

دمج SIMD مع الحلقات درس مجاني في Mojo Academy على CoddyKit. هذا هو الدرس 2 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Mojo Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Mojo Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

From Scalar to SIMD

To accelerate a kernel, swap the one-at-a-time loop for one that handles a whole pack of values per step with SIMD.

for i in range(n):
    out[i] = a[i] + b[i]

Pick a Width

Choose how many elements fit in one vector. The width is the number of lanes each step processes together.

alias width = 4

Load a Chunk

Grab several elements at once into a SIMD value with a vector load instead of reading them one by one.

var va = a.load[width=width](i)

Compute the Pack

Run the kernel's math on both chunks at once. The whole pack is added element-wise in a single operation.

var vsum = a.load[width=width](i) + b.load[width=width](i)

Store the Pack

Write the full result back with a vector store, covering every lane you just computed in one move.

out.store[width=width](i, vsum)

Step by the Width

The vectorized loop advances by the width, not by one. Each pass covers a full pack of elements.

for i in range(0, n, width):
    pass

Let vectorize Help

Mojo's vectorize helper sweeps a closure across the range in vector steps so you skip the manual bookkeeping.

from algorithm import vectorize

Write the Chunk Closure

You define a small parameterized function that handles one chunk. Mojo calls it with the right width as it sweeps.

fn body[w: Int](i: Int):
    out.store[width=w](i, a.load[width=w](i) + b.load[width=w](i))

Run vectorize

Call vectorize with your closure, the width, and the size. It loops and even handles the leftover tail for you.

vectorize[body, width](n)

Mind the Tail

When n is not a multiple of the width, a few elements remain. vectorize cleans up that tail so nothing is missed.

Same Output, More Speed

The vectorized kernel produces identical results but moves through data in big steps, so it finishes much sooner.

Quick Check

You vectorize a kernel with width 4 but n is 10. What handles the last two elements?

Recap

Vectorize a kernel by loading and storing packs, stepping by the width, and letting vectorize handle the sweep and the tail. 🚀

الأسئلة الشائعة

هل درس «دمج SIMD مع الحلقات» مجاني؟

نعم — نص درس «دمج SIMD مع الحلقات» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Mojo Academy، انتقل إلى CoddyKit PRO. تتضمن دورة Mojo Academy 4 دروس في المجموع.

ماذا ستتعلم في «دمج SIMD مع الحلقات»؟

وجّه جوهر النواة باستخدام المتجهات تتمرن على Mojo Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ Mojo Academy؟

لا تُشترط خبرة سابقة. Mojo Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 2 من أصل 4.

كم من الوقت يستغرق درس «دمج SIMD مع الحلقات»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس Mojo Academy هذا؟

نعم. كل درس في Mojo Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. تشريح نواة حسابية
  2. دمج SIMD مع الحلقات
  3. تقليل حركة الذاكرة
  4. التقسيم إلى كتل لتحسين محلية الذاكرة المخبئية
← العودة إلى Mojo Academy