Anatomie eines Compute-Kernels
Die heiße innere Schleife, die die Arbeit erledigt.
Anatomie eines Compute-Kernels ist eine kostenlose Mojo Academy-Lektion auf CoddyKit. Dies ist Lektion 1 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des Mojo Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der Mojo Academy-Kurs umfasst insgesamt 4 Lektionen.
Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.
What Is a Kernel?
A compute kernel is the small, focused routine that does the real numerical work, like adding two arrays element by element.
The Hot Inner Loop
Most of a kernel's time lives in one tight inner loop. Speed up that loop and you speed up the whole program.
for i in range(n):
out[i] = a[i] + b[i]Inputs and Outputs
A clean kernel takes its data buffers as parameters, so the same routine can run on any arrays you pass in.
fn add(a: UnsafePointer[Float32], b: UnsafePointer[Float32], out: UnsafePointer[Float32], n: Int):
passUse fn for Strictness
Kernels are written with fn, not def. The strict, typed style lets Mojo compile tight machine code with no surprises.
fn kernel(n: Int):
passKeep the Body Small
The fastest kernels do one thing. A short, predictable inner body is easy for the compiler to optimize aggressively.
No Surprises Inside
Avoid heavy work like allocation or I/O in the hot loop. Each iteration should be cheap and uniform for peak throughput.
Count the Work
Think in terms of total operations. A kernel over n elements does roughly n units of work, so n drives the cost.
Memory Is the Limit
Many kernels are not compute-bound but memory-bound. They wait on data, so how you move bytes often matters most.
Plan to Vectorize
Design the loop so each step can later process many elements with SIMD. A regular access pattern makes vectorizing easy.
A Tiny Saxpy Kernel
A classic kernel is saxpy: out = scale times a plus b. It is small, regular, and a great baseline to optimize.
for i in range(n):
out[i] = scale * a[i] + b[i]Measure Before You Tune
Start with the simple correct version and time it. That number is your baseline for every optimization that follows.
Quick Check
You are about to optimize a kernel. Where does almost all of its time go?
Recap
A kernel is a small fn whose tight inner loop does the work; keep it uniform, mind memory, and measure a baseline first. 🔧
Häufig gestellte Fragen
Ist die Lektion „Anatomie eines Compute-Kernels“ kostenlos?
Ja — der vollständige Text von „Anatomie eines Compute-Kernels“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des Mojo Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der Mojo Academy-Kurs umfasst insgesamt 4 Lektionen.
Was lerne ich in „Anatomie eines Compute-Kernels“?
Die heiße innere Schleife, die die Arbeit erledigt. Du übst Mojo Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.
Brauche ich Erfahrung, um Mojo Academy zu starten?
Keine Vorkenntnisse erforderlich. Mojo Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 1 von 4.
Wie lange dauert die Lektion „Anatomie eines Compute-Kernels“?
Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.
Kann ich in dieser Mojo Academy-Lektion Code schreiben und ausführen?
Ja. Jede Mojo Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.
Alle Lektionen in diesem Kurs
- Anatomie eines Compute-Kernels
- SIMD mit Schleifen kombinieren
- Speicherverkehr reduzieren
- Tiling für Cache-Lokalität