0Pricing
Mojo Academy · درس

دمج التوازي والمتجهات

أضف تعدد الخيوط فوق SIMD

دمج التوازي والمتجهات درس مجاني في Mojo Academy على CoddyKit. هذا هو الدرس 3 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Mojo Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Mojo Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

Two Kinds of Speed

SIMD packs many values into one core's instruction. Threads spread work across cores. Stacking both gives you the most. ⚡

Outer Layer: Threads

Use parallelize on the outer level so each core owns a chunk of the data. That is the coarse split across cores.

parallelize[do_chunk](workers)

Inner Layer: Vectors

Inside each chunk, use vectorize so one core chews through its slice in SIMD packs. That is the fine split per core.

from algorithm import vectorize

Import Both Helpers

Both tools live in the algorithm module, so you can bring them in together and combine them in one function.

from algorithm import parallelize, vectorize

The Chunk Worker

Each chunk worker computes its start and end, then hands its own slice to vectorize for the SIMD sweep.

fn do_chunk(c: Int):
    var start = c * chunk

Vectorize the Inner Range

Within the worker, call vectorize on just this chunk's length. The SIMD width sets how many lanes run at once.

vectorize[step, width](end - start)

Offset Into the Data

The inner step gets a local index, so add the chunk start to reach the right spot in the full array.

fn step[w: Int](i: Int):
    out.store[width=w](start + i, ...)

Pick the SIMD Width

Match the vector width to your hardware lanes with simdwidthof so each core uses its registers fully.

alias width = simdwidthof[DType.float32]()

Cores Times Lanes

The win multiplies: many cores, each doing many lanes per step. That product is why combined code is so fast.

Keep Chunks Independent

This stacking only works because chunks never touch each other's data. Independence keeps the two layers safe.

Measure the Combined Gain

Benchmark scalar, vector-only, and combined versions. The numbers reveal how much each layer contributes.

Quick Check

You combine threads and SIMD on the same workload.

Recap

You wrap parallelize over chunks and call vectorize inside each, offsetting by the chunk start, so cores times lanes multiply your throughput. 🚀

الأسئلة الشائعة

هل درس «دمج التوازي والمتجهات» مجاني؟

نعم — نص درس «دمج التوازي والمتجهات» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Mojo Academy، انتقل إلى CoddyKit PRO. تتضمن دورة Mojo Academy 4 دروس في المجموع.

ماذا ستتعلم في «دمج التوازي والمتجهات»؟

أضف تعدد الخيوط فوق SIMD تتمرن على Mojo Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ Mojo Academy؟

لا تُشترط خبرة سابقة. Mojo Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 3 من أصل 4.

كم من الوقت يستغرق درس «دمج التوازي والمتجهات»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس Mojo Academy هذا؟

نعم. كل درس في Mojo Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. دالة parallelize
  2. تقسيم العمل إلى أجزاء
  3. دمج التوازي والمتجهات
  4. تجنب سباقات البيانات
← العودة إلى Mojo Academy