Cooperative Groups
Una API moderna para ámbitos de sincronización flexibles.
Cooperative Groups es una lección gratuita de CUDA Academy en CoddyKit. Esta es la lección 4 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de CUDA Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de CUDA Academy incluye 4 lecciones en total.
Partes de esta lección aún no han sido traducidas y se muestran en inglés.
A Cleaner Sync API
Raw masks and intrinsics work, but they are fiddly. Cooperative groups wrap threads into objects you can name, size, and synchronize explicitly.
#include <cooperative_groups.h>
namespace cg = cooperative_groups;Grab the Whole Block
Start by getting a handle to your block of threads. this_thread_block returns a group you can sync just like syncthreads, but as an object.
cg::thread_block block = cg::this_thread_block();Sync Through the Group
Calling sync on the block is the modern barrier. It does exactly what syncthreads does, but reads clearly as a method on the group you mean.
block.sync();Carve Out a Warp Tile
You can split a block into fixed-size tiles. A 32-lane tile gives you a warp-sized group with clean methods instead of raw shuffle masks.
auto warp = cg::tiled_partition<32>(block);Methods Replace Masks
A tiled group offers shfl_down and friends without a mask argument. The group already knows its members, so the API stays short and safe.
val += warp.shfl_down(val, offset);Know Your Position
Every group exposes its members and your spot. thread_rank gives your index inside the group, and size returns how many threads it holds.
int rank = warp.thread_rank();Smaller Tiles Too
Tiles need not be 32 wide. A tiled_partition of 8 or 16 makes sub-warp groups, handy when your data naturally clusters in small sets.
Group-Level Reductions
The library ships ready-made collectives. A group reduce sums a tile in one call, hiding the offset loop you wrote by hand earlier.
int total = cg::reduce(warp, val, cg::plus<int>());Grids That Sync
The boldest group is the grid group. With a cooperative launch, every block can sync at one barrier, something a normal kernel cannot do.
Cooperative Launch Required
Grid-wide sync only works if you start the kernel with cudaLaunchCooperativeKernel and the GPU supports it. A normal launch will not allow it.
Why Bother
Cooperative groups make warp code readable and portable: no hand-managed masks, clear scopes, and reusable collectives that match what you mean.
Quick Check
Recall how you create a warp-sized cooperative group from a block.
Recap
Cooperative groups turn masks into named objects: tile a block, call reduce, even sync a whole grid. You now own warp-level CUDA. ✨
Preguntas frecuentes
¿La lección «Cooperative Groups» es gratis?
Sí — el texto completo de «Cooperative Groups» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de CUDA Academy, actualiza a CoddyKit PRO. El curso de CUDA Academy incluye 4 lecciones en total.
¿Qué aprenderé en «Cooperative Groups»?
Una API moderna para ámbitos de sincronización flexibles. Practicas CUDA Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.
¿Necesito experiencia previa para empezar CUDA Academy?
No se requiere experiencia previa. CUDA Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 4 de 4.
¿Cuánto tiempo toma la lección «Cooperative Groups»?
La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.
¿Puedo escribir y ejecutar código en esta lección de CUDA Academy?
Sí. Cada lección de CUDA Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.
Todas las lecciones de este curso
- Warps, lanes y máscaras
- __shfl_down_sync para reducciones
- Funciones de votación y ballot
- Cooperative Groups