0Pricing
CUDA Academy · Lektion

Ein mentales Modell der Hierarchie

Ordnen Sie Daten dem passenden Speicherbereich zu.

Ein mentales Modell der Hierarchie ist eine kostenlose CUDA Academy-Lektion auf CoddyKit. Dies ist Lektion 4 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des CUDA Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.

Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.

A Pyramid of Tradeoffs

Think of GPU memory as a pyramid. The top is tiny and lightning fast, while the base is huge but slow. Your job is to use each layer well.

Top: Registers

At the peak sit registers: per-thread, fastest, and very limited. Keep your hottest working values here whenever you possibly can.

Next: Shared Memory

Just below comes shared memory, an on-chip scratchpad visible to all threads in a block. It is your tool for fast intra-block teamwork.

Side Path: Constant Cache

Alongside sits the constant cache, ideal for small read-only values that a warp reads uniformly and broadcasts in one fetch.

Base: Global Memory

The wide base is global memory: gigabytes, visible to everyone, but high-latency. It holds your big inputs and outputs.

The Trap: Local Memory

Watch out for local memory. It sounds close by but lives in slow DRAM, used only when registers spill. Avoid relying on it.

Scope Decides the Space

Pick by who needs the data. One thread alone wants registers, a block working together wants shared memory, everyone wants global.

Lifetime Decides Too

Registers and shared memory vanish when a kernel ends, but global memory persists across launches. Match storage to how long data must live.

The Golden Rule

The winning pattern is load once from global, compute in fast on-chip memory, then write once back. This minimizes slow global memory traffic.

Capacity Costs Occupancy

Spending lots of registers or shared memory per block lets fewer blocks run at once. Occupancy is the balance you constantly tune.

Putting It Together

Great kernels deliberately route each piece of data to the right layer. That single habit is what separates slow code from fast CUDA code. 💪

Quick Check

Two threads in the SAME block need to share intermediate results. Which space fits best?

Recap: Match Data to Layer

You built a mental map: registers, shared, constant, and global trade speed for size and scope. Routing data to the right layer is the whole game. 🧠

Häufig gestellte Fragen

Ist die Lektion „Ein mentales Modell der Hierarchie“ kostenlos?

Ja — der vollständige Text von „Ein mentales Modell der Hierarchie“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des CUDA Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.

Was lerne ich in „Ein mentales Modell der Hierarchie“?

Ordnen Sie Daten dem passenden Speicherbereich zu. Du übst CUDA Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.

Brauche ich Erfahrung, um CUDA Academy zu starten?

Keine Vorkenntnisse erforderlich. CUDA Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 4 von 4.

Wie lange dauert die Lektion „Ein mentales Modell der Hierarchie“?

Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.

Kann ich in dieser CUDA Academy-Lektion Code schreiben und ausführen?

Ja. Jede CUDA Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.

Alle Lektionen in diesem Kurs

  1. Register und lokaler Speicher
  2. Kompromisse beim globalen Speicher
  3. Konstanter Speicher und sein Cache
  4. Ein mentales Modell der Hierarchie
← Zurück zu CUDA Academy