0Pricing
CUDA Academy · Урок

Создание гистограммы

Используйте атомарные операции с разделением данных в общей памяти.

«Создание гистограммы» — бесплатный урок CUDA Academy на CoddyKit. Это урок 3 из 4. Ты можешь прочитать весь урок бесплатно ниже — а потом практиковать его прямо в браузере с встроенным редактором кода и ИИ-репетитором 24/7. Это часть пути обучения CUDA Academy, и твой прогресс синхронизируется между веб-версией и приложением CoddyKit. Курс CUDA Academy содержит 4 уроков всего.

Части этого урока еще не переведены и отображаются на английском.

What a Histogram Counts

A histogram counts how many inputs fall into each bin. Many threads will want to increment the same bin, so this is an atomics problem. 📊

The Naive Approach

Each thread reads one element, finds its bin, and increments that bin. Without protection, popular bins lose counts to races.

The Global Atomic Version

The simplest correct fix is one atomicAdd per element straight into global memory. It works, but hot bins serialize threads.

atomicAdd(&hist[bin], 1);

The Contention Problem

When data clusters into a few bins, thousands of threads pile onto the same address. That contention can make global atomics painfully slow.

Privatization to the Rescue

Privatization gives each block its own private histogram in fast shared memory. Threads collide only inside their block, not across the whole grid.

Declare the Shared Histogram

Each block declares a shared array sized to the number of bins. It lives on-chip, so atomics there are far cheaper than global ones.

__shared__ int local[NBINS];

Step 1: Clear the Bins

Threads cooperatively zero the shared histogram, then call __syncthreads so no one counts before clearing is done.

local[tid] = 0;
__syncthreads();

Step 2: Count Locally

Now each thread atomically bumps its bin in shared memory. Same atomicAdd, but on the fast on-chip copy instead of global.

atomicAdd(&local[bin], 1);

Step 3: Merge to Global

After a barrier, threads add each shared bin into the global histogram with one atomicAdd per bin. Far fewer global atomics than before.

atomicAdd(&hist[i], local[i]);

Why This Is Faster

Shared-memory atomics are quick, and the costly global atomics now fire once per bin per block instead of once per element.

Watch the Bin Count

The private histogram must fit in shared memory. With too many bins, split them into passes or fall back to global atomics.

Quick Check

One question on the histogram strategy.

Recap: Building a Histogram

You built a histogram with global atomics, then sped it up using shared-memory privatization: clear, count locally, merge. ✅

Часто задаваемые вопросы

Урок «Создание гистограммы» бесплатный?

Да — полный текст урока «Создание гистограммы» бесплатно доступен здесь в веб-версии. Чтобы практиковать его интерактивно (встроенный редактор кода и ИИ-репетитор 24/7) и разблокировать остальной курс CUDA Academy, подпишись на CoddyKit PRO. Курс CUDA Academy содержит 4 уроков всего.

Чему я научусь в уроке «Создание гистограммы»?

Используйте атомарные операции с разделением данных в общей памяти. Ты практикуешь CUDA Academy с помощью реального кода, который запускаешь прямо в браузере, и ИИ-репетитор 24/7 отвечает на твои вопросы во время урока.

Нужен ли мне опыт, чтобы начать CUDA Academy?

Предыдущий опыт не требуется. CUDA Academy на CoddyKit структурирован для всех уровней — от новичков до продвинутых, поэтому ты можешь начать отсюда или с самого начала и учиться в своем темпе. Это урок 3 из 4.

Сколько времени занимает урок «Создание гистограммы»?

Большинство уроков CoddyKit занимают около 5–10 минут. Каждый из них компактный и интерактивный, поэтому ты постоянно делаешь прогресс и продолжаешь с того же места в веб-версии и приложении.

Можно ли писать и запускать код в этом уроке CUDA Academy?

Да. Каждый урок CUDA Academy включает встроенный редактор кода, поэтому ты пишешь и запускаешь реальный код прямо в браузере и получаешь моментальную обратную связь от AI — локальная установка не требуется.

Все уроки этого курса

  1. Состояния гонки на GPU
  2. atomicAdd и другие операции
  3. Создание гистограммы
  4. Пользовательские атомарные операции с atomicCAS
← Назад к CUDA Academy