0Pricing
CUDA Academy · درس

حلقات Grid-Stride

معالجة مصفوفات أكبر من الشبكة.

حلقات Grid-Stride درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

When the Array Is Huge

Sometimes the data is far bigger than the threads you launch. One thread per element no longer fits, so you need each thread to handle several elements.

The Grid Has a Size

The total number of threads is the grid size: blocks times threads per block. This stride is how far apart each thread's elements sit.

int stride = blockDim.x * gridDim.x;

Start, Then Stride

Each thread begins at its usual global index, then jumps forward by the grid size again and again until it runs off the array.

The Grid-Stride Loop

This loop is the whole pattern: start at i, step by stride, stop at n. Any array size is covered no matter how many threads you launch.

for (int i = blockIdx.x * blockDim.x + threadIdx.x;
     i < n;
     i += stride) {
    out[i] = a[i] + b[i];
}

The Built-in Guard

Notice the loop condition i < n is itself the bounds check. Threads that start past the end simply never enter the loop.

How the Work Splits

With a stride of 4 threads, thread 0 does elements 0, 4, 8, while thread 1 does 1, 5, 9. The work is interleaved, not chunked.

Interleaving Helps Coalescing

Because neighboring threads still touch neighboring addresses each step, the access stays coalesced and memory bandwidth stays high.

Decouple Threads From Data

Now your launch size no longer depends on n. You can pick a thread count that fits the GPU and let the loop absorb any workload.

Tune for the Hardware

A common choice is enough blocks to fill every multiprocessor, then let each thread loop. This keeps the GPU busy without overlaunching.

It Also Works When Tiny

If n is smaller than the grid, each thread runs the loop body at most once. The pattern degrades gracefully to the simple case.

A Robust Default

Many CUDA pros write every elementwise kernel as a grid-stride loop. It is flexible, safe, and rarely the wrong choice. 🚀

Quick Check

Identify the stride.

Recap

You learned the grid-stride loop: start at the global index and step by blockDim.x * gridDim.x until i reaches n. One kernel now handles any array size. 🎉

الأسئلة الشائعة

هل درس «حلقات Grid-Stride» مجاني؟

نعم — نص درس «حلقات Grid-Stride» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.

ماذا ستتعلم في «حلقات Grid-Stride»؟

معالجة مصفوفات أكبر من الشبكة. تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟

لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.

كم من الوقت يستغرق درس «حلقات Grid-Stride»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟

نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. صيغة الفهرس الكلاسيكية
  2. الحماية من الخروج عن النطاق
  3. التقريب إلى أعلى لعدد الكتل
  4. حلقات Grid-Stride
← العودة إلى CUDA Academy