SIMT: Dieselbe Instruktion, viele Threads
Das Ausführungsmodell, das GPUs schnell macht.
SIMT: Dieselbe Instruktion, viele Threads ist eine kostenlose CUDA Academy-Lektion auf CoddyKit. Dies ist Lektion 2 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des CUDA Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.
Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.
How the GPU Stays Busy
Thousands of cores need a smart way to be told what to do. The GPU's answer is SIMT: Single Instruction, Multiple Threads. ⚡
One Instruction, Many Threads
In SIMT, one instruction is broadcast to a whole group of threads at once. Each thread runs that same step, but on its own piece of data.
Same Recipe, Different Ingredients
Picture a kitchen where every cook follows the exact same recipe step, but each works on a different plate. That shared step is your instruction. 🍳
Meet the Warp
The GPU groups threads into bundles of 32 called a warp. A warp is the real unit that executes together, lockstep, one instruction at a time.
Why Bundles of 32
Issuing one instruction for 32 threads at once is far cheaper than 32 separate commands. That sharing is exactly where the GPU's efficiency comes from.
Each Thread Has Its Own Data
Threads in a warp share the instruction but keep private registers. So thread 0 and thread 5 run the same add, just on different numbers.
SIMT Is Not Quite SIMD
Classic SIMD processes fixed-width vectors. SIMT keeps the idea of shared instructions but lets each thread behave more independently when needed.
The Problem of Branches
What if half a warp takes an if branch and half does not? Threads in a warp want to march together, so a branch can split the group apart.
if (x > 0) {
y = x * 2;
} else {
y = -x;
}Warp Divergence
When threads in a warp disagree on a branch, the warp runs each path in turn and disables the others. This serial replay is called divergence.
Divergence Costs Speed
Because divergent paths run one after another, you lose parallelism. Keeping a warp on the same path is a key idea for fast kernels.
Why SIMT Scales So Well
With one instruction feeding 32 threads, and many warps in flight, the GPU keeps its math units packed. That is how SIMT turns into raw throughput.
Quick Check
Let us make sure the SIMT vocabulary is solid.
Recap: SIMT
SIMT broadcasts one instruction to a warp of 32 threads, each on its own data. Avoid divergent branches to keep every thread marching together. 👍
Häufig gestellte Fragen
Ist die Lektion „SIMT: Dieselbe Instruktion, viele Threads“ kostenlos?
Ja — der vollständige Text von „SIMT: Dieselbe Instruktion, viele Threads“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des CUDA Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.
Was lerne ich in „SIMT: Dieselbe Instruktion, viele Threads“?
Das Ausführungsmodell, das GPUs schnell macht. Du übst CUDA Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.
Brauche ich Erfahrung, um CUDA Academy zu starten?
Keine Vorkenntnisse erforderlich. CUDA Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 2 von 4.
Wie lange dauert die Lektion „SIMT: Dieselbe Instruktion, viele Threads“?
Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.
Kann ich in dieser CUDA Academy-Lektion Code schreiben und ausführen?
Ja. Jede CUDA Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.
Alle Lektionen in diesem Kurs
- CPU vs. GPU: Latenz vs. Durchsatz
- SIMT: Dieselbe Instruktion, viele Threads
- Was CUDA tatsächlich ist
- Probleme, die die GPU lieben