0Pricing
CUDA Academy · درس

صيغة الفهرس الكلاسيكية

blockIdx.x * blockDim.x + threadIdx.x.

صيغة الفهرس الكلاسيكية درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 1 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

One Thread, One Element

The whole point of a CUDA kernel is that every thread handles one piece of data. To do that, each thread needs a unique global index.

Why Local IDs Are Not Enough

Inside a block, threadIdx.x only counts 0 up to blockDim minus one. Many blocks reuse those same small numbers, so it cannot be your final index.

Blocks Sit Side by Side

Picture the grid as blocks laid end to end. Each block owns a contiguous slice of the array, and blockIdx.x tells you which slice you are in.

How Wide Is a Block?

blockDim.x is the number of threads per block. It is the width of each slice, so it scales your block offset to the right spot.

The Classic Formula

Combine the three: skip past earlier blocks, then add your position inside this block. That single line gives every thread a unique index.

int i = blockIdx.x * blockDim.x + threadIdx.x;

Walking Through It

With 4 threads per block, block 0 covers 0 to 3, block 1 covers 4 to 7, block 2 covers 8 to 11. The offsets never overlap.

A Concrete Example

Thread 2 in block 3 with blockDim 256 lands at 3 times 256 plus 2, which is 770. That is its global position in the data.

// blockIdx.x=3, blockDim.x=256, threadIdx.x=2
int i = 3 * 256 + 2; // i == 770

Using the Index

Once you have i, you treat it as an array subscript. Each thread reads and writes only its own element, with no overlap.

out[i] = a[i] + b[i];

It Maps One to One

Launch n threads and the formula produces every value from 0 to n minus 1 exactly once. That is a perfect one-to-one cover of the array.

The Order Matters

Always multiply before you add. blockIdx.x * blockDim.x is the start of your slice, and threadIdx.x is the step inside it.

Beyond One Dimension

The same idea extends to 2D and 3D using the .y and .z members, but for flat arrays the .x formula is all you need. 🚀

Quick Check

Compute one thread's global index.

Recap

You learned the formula every kernel uses: blockIdx.x * blockDim.x + threadIdx.x. It hands each thread one unique slot in your array. 🎉

الأسئلة الشائعة

هل درس «صيغة الفهرس الكلاسيكية» مجاني؟

نعم — نص درس «صيغة الفهرس الكلاسيكية» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.

ماذا ستتعلم في «صيغة الفهرس الكلاسيكية»؟

blockIdx.x * blockDim.x + threadIdx.x. تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟

لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 1 من أصل 4.

كم من الوقت يستغرق درس «صيغة الفهرس الكلاسيكية»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟

نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. صيغة الفهرس الكلاسيكية
  2. الحماية من الخروج عن النطاق
  3. التقريب إلى أعلى لعدد الكتل
  4. حلقات Grid-Stride
← العودة إلى CUDA Academy