Lecturas fusionadas frente a estriadas
Descubra cómo importa la relación entre hilos y direcciones.
Lecturas fusionadas frente a estriadas es una lección gratuita de CUDA Academy en CoddyKit. Esta es la lección 2 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de CUDA Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de CUDA Academy incluye 4 lecciones en total.
Partes de esta lección aún no han sido traducidas y se muestran en inglés.
It Is All About Mapping
Coalescing depends on which address each thread touches. The map from thread to address decides whether one transaction serves the whole warp.
The Coalesced Pattern
When neighboring threads read neighboring addresses, the warp covers one contiguous run. That tidy layout is a coalesced access.
int i = blockIdx.x * blockDim.x + threadIdx.x;
float v = data[i];Why It Wins
Thread 0 hits index 0, thread 1 hits index 1, and so on. All 32 addresses fall in one aligned line, so the GPU needs a single transaction.
Enter the Stride
A stride is a fixed gap between the addresses successive threads touch. The moment that gap grows past one element, coalescing starts to break.
float v = data[i * stride];Strided Reads Spread Out
With a stride of 2, thread 0 reads 0, thread 1 reads 2, thread 2 reads 4. The warp now spans twice the memory, so it needs more transactions.
The Cost Scales Up
Double the stride and you roughly double the lines touched. Big strides can push a warp toward many transactions while using only a sliver of each line.
A Common Trap
Giving each thread a whole column of a row-major matrix creates a huge stride. It looks neat in code but quietly scatters every warp.
float v = matrix[threadIdx.x * width + row];Transpose the Mapping
Often the fix is just swapping which index varies fastest with the thread. Let consecutive threads walk consecutive memory and the reads coalesce again.
float v = matrix[row * width + threadIdx.x];Offsets Hurt Too
Even a coalesced pattern suffers if it starts mid-line. A misaligned offset shifts the warp across a boundary, splitting one read into two.
Think in Warps
To judge a pattern, do not picture one thread. Picture all 32 lanes at once and ask how many lines their addresses cover together. 🔍
The Rule of Thumb
Make the fastest-changing index follow threadIdx.x. That single habit keeps most of your global reads coalesced for free.
Quick Check
Pick the access pattern that coalesces best.
Recap
You saw that coalesced reads keep neighbors together while a stride scatters them across lines. Keep threadIdx.x driving the fastest index. 🎉
Preguntas frecuentes
¿La lección «Lecturas fusionadas frente a estriadas» es gratis?
Sí — el texto completo de «Lecturas fusionadas frente a estriadas» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de CUDA Academy, actualiza a CoddyKit PRO. El curso de CUDA Academy incluye 4 lecciones en total.
¿Qué aprenderé en «Lecturas fusionadas frente a estriadas»?
Descubra cómo importa la relación entre hilos y direcciones. Practicas CUDA Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.
¿Necesito experiencia previa para empezar CUDA Academy?
No se requiere experiencia previa. CUDA Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 2 de 4.
¿Cuánto tiempo toma la lección «Lecturas fusionadas frente a estriadas»?
La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.
¿Puedo escribir y ejecutar código en esta lección de CUDA Academy?
Sí. Cada lección de CUDA Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.
Todas las lecciones de este curso
- Qué es una transacción de memoria
- Lecturas fusionadas frente a estriadas
- Estructura de arrays frente a array de estructuras
- Medir el ancho de banda efectivo