دمج التوازي والمتجهات
أضف تعدد الخيوط فوق SIMD
دمج التوازي والمتجهات درس مجاني في Mojo Academy على CoddyKit. هذا هو الدرس 3 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Mojo Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Mojo Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
Two Kinds of Speed
SIMD packs many values into one core's instruction. Threads spread work across cores. Stacking both gives you the most. ⚡
Outer Layer: Threads
Use parallelize on the outer level so each core owns a chunk of the data. That is the coarse split across cores.
parallelize[do_chunk](workers)Inner Layer: Vectors
Inside each chunk, use vectorize so one core chews through its slice in SIMD packs. That is the fine split per core.
from algorithm import vectorizeImport Both Helpers
Both tools live in the algorithm module, so you can bring them in together and combine them in one function.
from algorithm import parallelize, vectorizeThe Chunk Worker
Each chunk worker computes its start and end, then hands its own slice to vectorize for the SIMD sweep.
fn do_chunk(c: Int):
var start = c * chunkVectorize the Inner Range
Within the worker, call vectorize on just this chunk's length. The SIMD width sets how many lanes run at once.
vectorize[step, width](end - start)Offset Into the Data
The inner step gets a local index, so add the chunk start to reach the right spot in the full array.
fn step[w: Int](i: Int):
out.store[width=w](start + i, ...)Pick the SIMD Width
Match the vector width to your hardware lanes with simdwidthof so each core uses its registers fully.
alias width = simdwidthof[DType.float32]()Cores Times Lanes
The win multiplies: many cores, each doing many lanes per step. That product is why combined code is so fast.
Keep Chunks Independent
This stacking only works because chunks never touch each other's data. Independence keeps the two layers safe.
Measure the Combined Gain
Benchmark scalar, vector-only, and combined versions. The numbers reveal how much each layer contributes.
Quick Check
You combine threads and SIMD on the same workload.
Recap
You wrap parallelize over chunks and call vectorize inside each, offsetting by the chunk start, so cores times lanes multiply your throughput. 🚀
الأسئلة الشائعة
هل درس «دمج التوازي والمتجهات» مجاني؟
نعم — نص درس «دمج التوازي والمتجهات» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Mojo Academy، انتقل إلى CoddyKit PRO. تتضمن دورة Mojo Academy 4 دروس في المجموع.
ماذا ستتعلم في «دمج التوازي والمتجهات»؟
أضف تعدد الخيوط فوق SIMD تتمرن على Mojo Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ Mojo Academy؟
لا تُشترط خبرة سابقة. Mojo Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 3 من أصل 4.
كم من الوقت يستغرق درس «دمج التوازي والمتجهات»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس Mojo Academy هذا؟
نعم. كل درس في Mojo Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- دالة parallelize
- تقسيم العمل إلى أجزاء
- دمج التوازي والمتجهات
- تجنب سباقات البيانات