Konstanter Speicher und sein Cache
Lesen Sie schreibgeschützte Werte kostengünstig per Broadcast aus.
Konstanter Speicher und sein Cache ist eine kostenlose CUDA Academy-Lektion auf CoddyKit. Dies ist Lektion 3 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des CUDA Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.
Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.
A Space for Read-Only Values
Constant memory is a small region built for data that every thread reads but none of them ever writes during a kernel.
Declaring It
You mark a global-scope array with the __constant__ qualifier. It lives outside any function, visible to all your kernels.
__constant__ float weights[256];It Is Genuinely Small
Total constant memory is only 64 KB on typical GPUs. It is meant for coefficients and parameters, not for large datasets.
Filling It from the Host
You write to constant memory from the CPU using cudaMemcpyToSymbol, since kernels themselves cannot modify it.
cudaMemcpyToSymbol(weights, h_w,
256 * sizeof(float));Backed by a Special Cache
Reads flow through a dedicated constant cache on each multiprocessor. A cached value comes back almost as fast as a register.
The Broadcast Superpower
When a whole warp reads the same address, the hardware does one fetch and broadcasts it to all 32 threads in a single step. 📡
Same Address Is the Key
The win depends on uniform access. If threads in a warp read different constant addresses, those reads serialize and the speedup is lost.
Reading It Looks Normal
Inside a kernel you read constant memory like any array. The compiler quietly routes the access through the fast constant cache.
__global__ void k(float *out, int i) {
out[i] = weights[0] * 2.0f;
}A Perfect Fit for Filters
Convolution kernels and lookup tables shine here, because every thread reuses the same small set of coefficients over and over.
Set Once, Reuse Often
You typically upload to constant memory a single time, then launch many kernels that all read those unchanging values cheaply.
When Not to Use It
Skip constant memory if your data is large or if threads read scattered addresses. In those cases plain global memory serves you better.
Quick Check
When does constant memory give its biggest speedup?
Recap: Small, Cached, Broadcast
You learned constant memory is a tiny read-only space whose cache broadcasts uniform reads to a whole warp. Use it for small, shared values. ✨
Häufig gestellte Fragen
Ist die Lektion „Konstanter Speicher und sein Cache“ kostenlos?
Ja — der vollständige Text von „Konstanter Speicher und sein Cache“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des CUDA Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.
Was lerne ich in „Konstanter Speicher und sein Cache“?
Lesen Sie schreibgeschützte Werte kostengünstig per Broadcast aus. Du übst CUDA Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.
Brauche ich Erfahrung, um CUDA Academy zu starten?
Keine Vorkenntnisse erforderlich. CUDA Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 3 von 4.
Wie lange dauert die Lektion „Konstanter Speicher und sein Cache“?
Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.
Kann ich in dieser CUDA Academy-Lektion Code schreiben und ausführen?
Ja. Jede CUDA Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.
Alle Lektionen in diesem Kurs
- Register und lokaler Speicher
- Kompromisse beim globalen Speicher
- Konstanter Speicher und sein Cache
- Ein mentales Modell der Hierarchie