0Pricing
CUDA Academy · درس

Cooperative Groups

واجهة API حديثة لنطاقات مزامنة مرنة.

Cooperative Groups درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

A Cleaner Sync API

Raw masks and intrinsics work, but they are fiddly. Cooperative groups wrap threads into objects you can name, size, and synchronize explicitly.

#include <cooperative_groups.h>
namespace cg = cooperative_groups;

Grab the Whole Block

Start by getting a handle to your block of threads. this_thread_block returns a group you can sync just like syncthreads, but as an object.

cg::thread_block block = cg::this_thread_block();

Sync Through the Group

Calling sync on the block is the modern barrier. It does exactly what syncthreads does, but reads clearly as a method on the group you mean.

block.sync();

Carve Out a Warp Tile

You can split a block into fixed-size tiles. A 32-lane tile gives you a warp-sized group with clean methods instead of raw shuffle masks.

auto warp = cg::tiled_partition<32>(block);

Methods Replace Masks

A tiled group offers shfl_down and friends without a mask argument. The group already knows its members, so the API stays short and safe.

val += warp.shfl_down(val, offset);

Know Your Position

Every group exposes its members and your spot. thread_rank gives your index inside the group, and size returns how many threads it holds.

int rank = warp.thread_rank();

Smaller Tiles Too

Tiles need not be 32 wide. A tiled_partition of 8 or 16 makes sub-warp groups, handy when your data naturally clusters in small sets.

Group-Level Reductions

The library ships ready-made collectives. A group reduce sums a tile in one call, hiding the offset loop you wrote by hand earlier.

int total = cg::reduce(warp, val, cg::plus<int>());

Grids That Sync

The boldest group is the grid group. With a cooperative launch, every block can sync at one barrier, something a normal kernel cannot do.

Cooperative Launch Required

Grid-wide sync only works if you start the kernel with cudaLaunchCooperativeKernel and the GPU supports it. A normal launch will not allow it.

Why Bother

Cooperative groups make warp code readable and portable: no hand-managed masks, clear scopes, and reusable collectives that match what you mean.

Quick Check

Recall how you create a warp-sized cooperative group from a block.

Recap

Cooperative groups turn masks into named objects: tile a block, call reduce, even sync a whole grid. You now own warp-level CUDA. ✨

الأسئلة الشائعة

هل درس «Cooperative Groups» مجاني؟

نعم — نص درس «Cooperative Groups» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.

ماذا ستتعلم في «Cooperative Groups»؟

واجهة API حديثة لنطاقات مزامنة مرنة. تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟

لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.

كم من الوقت يستغرق درس «Cooperative Groups»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟

نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. Warps وLanes وMasks
  2. استخدام __shfl_down_sync للاختزال
  3. دوال Ballot وVote
  4. Cooperative Groups
← العودة إلى CUDA Academy