دمج SIMD مع الحلقات
وجّه جوهر النواة باستخدام المتجهات
دمج SIMD مع الحلقات درس مجاني في Mojo Academy على CoddyKit. هذا هو الدرس 2 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Mojo Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Mojo Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
From Scalar to SIMD
To accelerate a kernel, swap the one-at-a-time loop for one that handles a whole pack of values per step with SIMD.
for i in range(n):
out[i] = a[i] + b[i]Pick a Width
Choose how many elements fit in one vector. The width is the number of lanes each step processes together.
alias width = 4Load a Chunk
Grab several elements at once into a SIMD value with a vector load instead of reading them one by one.
var va = a.load[width=width](i)Compute the Pack
Run the kernel's math on both chunks at once. The whole pack is added element-wise in a single operation.
var vsum = a.load[width=width](i) + b.load[width=width](i)Store the Pack
Write the full result back with a vector store, covering every lane you just computed in one move.
out.store[width=width](i, vsum)Step by the Width
The vectorized loop advances by the width, not by one. Each pass covers a full pack of elements.
for i in range(0, n, width):
passLet vectorize Help
Mojo's vectorize helper sweeps a closure across the range in vector steps so you skip the manual bookkeeping.
from algorithm import vectorizeWrite the Chunk Closure
You define a small parameterized function that handles one chunk. Mojo calls it with the right width as it sweeps.
fn body[w: Int](i: Int):
out.store[width=w](i, a.load[width=w](i) + b.load[width=w](i))Run vectorize
Call vectorize with your closure, the width, and the size. It loops and even handles the leftover tail for you.
vectorize[body, width](n)Mind the Tail
When n is not a multiple of the width, a few elements remain. vectorize cleans up that tail so nothing is missed.
Same Output, More Speed
The vectorized kernel produces identical results but moves through data in big steps, so it finishes much sooner.
Quick Check
You vectorize a kernel with width 4 but n is 10. What handles the last two elements?
Recap
Vectorize a kernel by loading and storing packs, stepping by the width, and letting vectorize handle the sweep and the tail. 🚀
الأسئلة الشائعة
هل درس «دمج SIMD مع الحلقات» مجاني؟
نعم — نص درس «دمج SIMD مع الحلقات» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Mojo Academy، انتقل إلى CoddyKit PRO. تتضمن دورة Mojo Academy 4 دروس في المجموع.
ماذا ستتعلم في «دمج SIMD مع الحلقات»؟
وجّه جوهر النواة باستخدام المتجهات تتمرن على Mojo Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ Mojo Academy؟
لا تُشترط خبرة سابقة. Mojo Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 2 من أصل 4.
كم من الوقت يستغرق درس «دمج SIMD مع الحلقات»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس Mojo Academy هذا؟
نعم. كل درس في Mojo Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.