0Pricing
CUDA Academy · Lección

Declarar arrays __shared__

Memoria scratchpad rápida por bloque.

Declarar arrays __shared__ es una lección gratuita de CUDA Academy en CoddyKit. Esta es la lección 1 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de CUDA Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de CUDA Academy incluye 4 lecciones en total.

Partes de esta lección aún no han sido traducidas y se muestran en inglés.

A Scratchpad Per Block

Every block gets a tiny, fast pool of on-chip memory called shared memory. Think of it as a scratchpad the whole block can write to and read from together.

Why It Is So Fast

Shared memory sits right on the streaming multiprocessor, so it is roughly 100x faster than global memory. Use it to avoid hammering slow off-chip DRAM. 🚀

The __shared__ Keyword

You declare it with the __shared__ qualifier inside a kernel. This one array is then shared by every thread in the block.

__global__ void kern() {
    __shared__ float tile[256];
}

Fixed Size at Compile Time

When you give a size in brackets, that is static shared memory. The compiler must know the size, so it has to be a constant, not a runtime value.

__shared__ int counts[128];

One Copy, Not One Per Thread

This is the key idea: a __shared__ array is created once per block, not once per thread. All threads see the exact same array.

Threads Cooperate Through It

Because every thread sees the same data, shared memory lets threads cooperate. One thread can stash a value and a neighbor can pick it up.

Block Scope, Not Beyond

Its lifetime matches the block. The array is born when the block starts and gone when it ends, so it has block scope only. 🧱

Each Thread Owns a Slot

A common pattern is one slot per thread, indexed by threadIdx.x. Each thread loads its element into shared memory in parallel.

__shared__ float s[256];
s[threadIdx.x] = input[i];

A Tiny but Precious Resource

Shared memory is small, often just 48 to 100 KB per SM. Asking for too much per block reduces how many blocks can run at once.

The Classic Use: Staging Tiles

The most common job is staging a tile of global data so the block can reuse it many times without going back to slow DRAM.

Not Visible to Other Blocks

Remember the boundary: shared memory is private to its block. Two different blocks each get their own separate copy and cannot peek at each other.

Quick Check

Let us check how shared memory is scoped.

Recap

You learned that __shared__ gives each block a fast on-chip scratchpad, created once per block and ideal for staging reusable data. Next: keeping threads in step. 🎯

Preguntas frecuentes

¿La lección «Declarar arrays __shared__» es gratis?

Sí — el texto completo de «Declarar arrays __shared__» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de CUDA Academy, actualiza a CoddyKit PRO. El curso de CUDA Academy incluye 4 lecciones en total.

¿Qué aprenderé en «Declarar arrays __shared__»?

Memoria scratchpad rápida por bloque. Practicas CUDA Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.

¿Necesito experiencia previa para empezar CUDA Academy?

No se requiere experiencia previa. CUDA Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 1 de 4.

¿Cuánto tiempo toma la lección «Declarar arrays __shared__»?

La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.

¿Puedo escribir y ejecutar código en esta lección de CUDA Academy?

Sí. Cada lección de CUDA Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.

Todas las lecciones de este curso

  1. Declarar arrays __shared__
  2. Sincronizar con __syncthreads
  3. Evitar conflictos de bancos
  4. Memoria compartida dinámica
← Volver a CUDA Academy