0Pricing
CUDA Academy · Lección

Construir un histograma

Use atómicos con privatización en memoria compartida.

Construir un histograma es una lección gratuita de CUDA Academy en CoddyKit. Esta es la lección 3 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de CUDA Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de CUDA Academy incluye 4 lecciones en total.

Partes de esta lección aún no han sido traducidas y se muestran en inglés.

What a Histogram Counts

A histogram counts how many inputs fall into each bin. Many threads will want to increment the same bin, so this is an atomics problem. 📊

The Naive Approach

Each thread reads one element, finds its bin, and increments that bin. Without protection, popular bins lose counts to races.

The Global Atomic Version

The simplest correct fix is one atomicAdd per element straight into global memory. It works, but hot bins serialize threads.

atomicAdd(&hist[bin], 1);

The Contention Problem

When data clusters into a few bins, thousands of threads pile onto the same address. That contention can make global atomics painfully slow.

Privatization to the Rescue

Privatization gives each block its own private histogram in fast shared memory. Threads collide only inside their block, not across the whole grid.

Declare the Shared Histogram

Each block declares a shared array sized to the number of bins. It lives on-chip, so atomics there are far cheaper than global ones.

__shared__ int local[NBINS];

Step 1: Clear the Bins

Threads cooperatively zero the shared histogram, then call __syncthreads so no one counts before clearing is done.

local[tid] = 0;
__syncthreads();

Step 2: Count Locally

Now each thread atomically bumps its bin in shared memory. Same atomicAdd, but on the fast on-chip copy instead of global.

atomicAdd(&local[bin], 1);

Step 3: Merge to Global

After a barrier, threads add each shared bin into the global histogram with one atomicAdd per bin. Far fewer global atomics than before.

atomicAdd(&hist[i], local[i]);

Why This Is Faster

Shared-memory atomics are quick, and the costly global atomics now fire once per bin per block instead of once per element.

Watch the Bin Count

The private histogram must fit in shared memory. With too many bins, split them into passes or fall back to global atomics.

Quick Check

One question on the histogram strategy.

Recap: Building a Histogram

You built a histogram with global atomics, then sped it up using shared-memory privatization: clear, count locally, merge. ✅

Preguntas frecuentes

¿La lección «Construir un histograma» es gratis?

Sí — el texto completo de «Construir un histograma» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de CUDA Academy, actualiza a CoddyKit PRO. El curso de CUDA Academy incluye 4 lecciones en total.

¿Qué aprenderé en «Construir un histograma»?

Use atómicos con privatización en memoria compartida. Practicas CUDA Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.

¿Necesito experiencia previa para empezar CUDA Academy?

No se requiere experiencia previa. CUDA Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 3 de 4.

¿Cuánto tiempo toma la lección «Construir un histograma»?

La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.

¿Puedo escribir y ejecutar código en esta lección de CUDA Academy?

Sí. Cada lección de CUDA Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.

Todas las lecciones de este curso

  1. Condiciones de carrera en la GPU
  2. atomicAdd y funciones similares
  3. Construir un histograma
  4. Atómicos personalizados con atomicCAS
← Volver a CUDA Academy