Grid-Stride-Schleifen
Verarbeiten Sie Arrays, die größer als das Grid sind.
Grid-Stride-Schleifen ist eine kostenlose CUDA Academy-Lektion auf CoddyKit. Dies ist Lektion 4 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des CUDA Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.
Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.
When the Array Is Huge
Sometimes the data is far bigger than the threads you launch. One thread per element no longer fits, so you need each thread to handle several elements.
The Grid Has a Size
The total number of threads is the grid size: blocks times threads per block. This stride is how far apart each thread's elements sit.
int stride = blockDim.x * gridDim.x;Start, Then Stride
Each thread begins at its usual global index, then jumps forward by the grid size again and again until it runs off the array.
The Grid-Stride Loop
This loop is the whole pattern: start at i, step by stride, stop at n. Any array size is covered no matter how many threads you launch.
for (int i = blockIdx.x * blockDim.x + threadIdx.x;
i < n;
i += stride) {
out[i] = a[i] + b[i];
}The Built-in Guard
Notice the loop condition i < n is itself the bounds check. Threads that start past the end simply never enter the loop.
How the Work Splits
With a stride of 4 threads, thread 0 does elements 0, 4, 8, while thread 1 does 1, 5, 9. The work is interleaved, not chunked.
Interleaving Helps Coalescing
Because neighboring threads still touch neighboring addresses each step, the access stays coalesced and memory bandwidth stays high.
Decouple Threads From Data
Now your launch size no longer depends on n. You can pick a thread count that fits the GPU and let the loop absorb any workload.
Tune for the Hardware
A common choice is enough blocks to fill every multiprocessor, then let each thread loop. This keeps the GPU busy without overlaunching.
It Also Works When Tiny
If n is smaller than the grid, each thread runs the loop body at most once. The pattern degrades gracefully to the simple case.
A Robust Default
Many CUDA pros write every elementwise kernel as a grid-stride loop. It is flexible, safe, and rarely the wrong choice. 🚀
Quick Check
Identify the stride.
Recap
You learned the grid-stride loop: start at the global index and step by blockDim.x * gridDim.x until i reaches n. One kernel now handles any array size. 🎉
Häufig gestellte Fragen
Ist die Lektion „Grid-Stride-Schleifen“ kostenlos?
Ja — der vollständige Text von „Grid-Stride-Schleifen“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des CUDA Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.
Was lerne ich in „Grid-Stride-Schleifen“?
Verarbeiten Sie Arrays, die größer als das Grid sind. Du übst CUDA Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.
Brauche ich Erfahrung, um CUDA Academy zu starten?
Keine Vorkenntnisse erforderlich. CUDA Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 4 von 4.
Wie lange dauert die Lektion „Grid-Stride-Schleifen“?
Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.
Kann ich in dieser CUDA Academy-Lektion Code schreiben und ausführen?
Ja. Jede CUDA Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.
Alle Lektionen in diesem Kurs
- Die klassische Indexformel
- Zugriffe außerhalb des Bereichs verhindern
- Die Blockanzahl aufrunden
- Grid-Stride-Schleifen