0Pricing
CUDA Academy · درس

بناء Histogram

استخدم العمليات الذرية مع تخصيص الذاكرة المشتركة لكل جزء.

بناء Histogram درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 3 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

What a Histogram Counts

A histogram counts how many inputs fall into each bin. Many threads will want to increment the same bin, so this is an atomics problem. 📊

The Naive Approach

Each thread reads one element, finds its bin, and increments that bin. Without protection, popular bins lose counts to races.

The Global Atomic Version

The simplest correct fix is one atomicAdd per element straight into global memory. It works, but hot bins serialize threads.

atomicAdd(&hist[bin], 1);

The Contention Problem

When data clusters into a few bins, thousands of threads pile onto the same address. That contention can make global atomics painfully slow.

Privatization to the Rescue

Privatization gives each block its own private histogram in fast shared memory. Threads collide only inside their block, not across the whole grid.

Declare the Shared Histogram

Each block declares a shared array sized to the number of bins. It lives on-chip, so atomics there are far cheaper than global ones.

__shared__ int local[NBINS];

Step 1: Clear the Bins

Threads cooperatively zero the shared histogram, then call __syncthreads so no one counts before clearing is done.

local[tid] = 0;
__syncthreads();

Step 2: Count Locally

Now each thread atomically bumps its bin in shared memory. Same atomicAdd, but on the fast on-chip copy instead of global.

atomicAdd(&local[bin], 1);

Step 3: Merge to Global

After a barrier, threads add each shared bin into the global histogram with one atomicAdd per bin. Far fewer global atomics than before.

atomicAdd(&hist[i], local[i]);

Why This Is Faster

Shared-memory atomics are quick, and the costly global atomics now fire once per bin per block instead of once per element.

Watch the Bin Count

The private histogram must fit in shared memory. With too many bins, split them into passes or fall back to global atomics.

Quick Check

One question on the histogram strategy.

Recap: Building a Histogram

You built a histogram with global atomics, then sped it up using shared-memory privatization: clear, count locally, merge. ✅

الأسئلة الشائعة

هل درس «بناء Histogram» مجاني؟

نعم — نص درس «بناء Histogram» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.

ماذا ستتعلم في «بناء Histogram»؟

استخدم العمليات الذرية مع تخصيص الذاكرة المشتركة لكل جزء. تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟

لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 3 من أصل 4.

كم من الوقت يستغرق درس «بناء Histogram»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟

نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. حالات التسابق على GPU
  2. atomicAdd وما شابهها
  3. بناء Histogram
  4. عمليات ذرية مخصصة باستخدام atomicCAS
← العودة إلى CUDA Academy