Warps وLanes وMasks
وحدة الخيوط ذات الـ32 خيطًا والأقنعة النشطة.
Warps وLanes وMasks درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 1 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
Threads Run in Warps
The GPU does not schedule threads one by one. It groups them into a warp of 32 threads that march together, executing the same instruction in lockstep.
Why 32 Matters
On NVIDIA hardware a warp is always 32 threads. It is the real unit of execution, so good kernels think in groups of 32, not single threads.
Each Thread Is a Lane
Inside a warp, every thread has a position from 0 to 31 called its lane. The lane id is how warp primitives know who is talking to whom.
int lane = threadIdx.x % 32;Finding the Lane Id
You can read the lane directly from a special register instead of computing it. The laneid register always holds 0 through 31 for the current thread.
Lockstep Has a Catch
Threads share one program counter per warp. When an if sends lanes down different paths, the warp diverges and runs each path in turn, hurting speed.
The Active Mask
Not every lane is always running. A 32-bit mask marks which lanes are active right now, with one bit per lane set to 1 when that lane participates.
Why Masks Exist
After divergence only some lanes are live. Warp primitives need to know exactly who is present, so you pass them an active mask to stay correct.
The Full Mask
When you are sure all 32 lanes are active, the mask is 0xffffffff, every bit set. This full mask is the most common value you will pass.
unsigned mask = 0xffffffff;Building a Mask Safely
Inside a branch, do not guess the mask. Call activemask to capture exactly which lanes reached this point right now.
unsigned mask = __activemask();Lanes Talk Without Memory
The big win is that lanes in one warp can swap data directly through registers. No shared memory and no barriers are needed for warp-local exchange.
Sync Means the Warp
The sync suffix on these intrinsics is a warp-level handshake, not a block barrier. It only coordinates the lanes named in the mask you give it.
Quick Check
Recall what a warp is and how big it is on NVIDIA GPUs.
Recap
A warp is 32 lanes running in lockstep, and a mask tracks who is active. That foundation lets lanes share data fast. Next: shuffles for reductions. ✨
الأسئلة الشائعة
هل درس «Warps وLanes وMasks» مجاني؟
نعم — نص درس «Warps وLanes وMasks» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.
ماذا ستتعلم في «Warps وLanes وMasks»؟
وحدة الخيوط ذات الـ32 خيطًا والأقنعة النشطة. تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟
لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 1 من أصل 4.
كم من الوقت يستغرق درس «Warps وLanes وMasks»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟
نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- Warps وLanes وMasks
- استخدام __shfl_down_sync للاختزال
- دوال Ballot وVote
- Cooperative Groups