0Pricing
CUDA Academy · Lektion

Warps, Lanes und Masken

Die Einheit aus 32 Threads und aktive Masken.

Warps, Lanes und Masken ist eine kostenlose CUDA Academy-Lektion auf CoddyKit. Dies ist Lektion 1 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des CUDA Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.

Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.

Threads Run in Warps

The GPU does not schedule threads one by one. It groups them into a warp of 32 threads that march together, executing the same instruction in lockstep.

Why 32 Matters

On NVIDIA hardware a warp is always 32 threads. It is the real unit of execution, so good kernels think in groups of 32, not single threads.

Each Thread Is a Lane

Inside a warp, every thread has a position from 0 to 31 called its lane. The lane id is how warp primitives know who is talking to whom.

int lane = threadIdx.x % 32;

Finding the Lane Id

You can read the lane directly from a special register instead of computing it. The laneid register always holds 0 through 31 for the current thread.

Lockstep Has a Catch

Threads share one program counter per warp. When an if sends lanes down different paths, the warp diverges and runs each path in turn, hurting speed.

The Active Mask

Not every lane is always running. A 32-bit mask marks which lanes are active right now, with one bit per lane set to 1 when that lane participates.

Why Masks Exist

After divergence only some lanes are live. Warp primitives need to know exactly who is present, so you pass them an active mask to stay correct.

The Full Mask

When you are sure all 32 lanes are active, the mask is 0xffffffff, every bit set. This full mask is the most common value you will pass.

unsigned mask = 0xffffffff;

Building a Mask Safely

Inside a branch, do not guess the mask. Call activemask to capture exactly which lanes reached this point right now.

unsigned mask = __activemask();

Lanes Talk Without Memory

The big win is that lanes in one warp can swap data directly through registers. No shared memory and no barriers are needed for warp-local exchange.

Sync Means the Warp

The sync suffix on these intrinsics is a warp-level handshake, not a block barrier. It only coordinates the lanes named in the mask you give it.

Quick Check

Recall what a warp is and how big it is on NVIDIA GPUs.

Recap

A warp is 32 lanes running in lockstep, and a mask tracks who is active. That foundation lets lanes share data fast. Next: shuffles for reductions. ✨

Häufig gestellte Fragen

Ist die Lektion „Warps, Lanes und Masken“ kostenlos?

Ja — der vollständige Text von „Warps, Lanes und Masken“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des CUDA Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.

Was lerne ich in „Warps, Lanes und Masken“?

Die Einheit aus 32 Threads und aktive Masken. Du übst CUDA Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.

Brauche ich Erfahrung, um CUDA Academy zu starten?

Keine Vorkenntnisse erforderlich. CUDA Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 1 von 4.

Wie lange dauert die Lektion „Warps, Lanes und Masken“?

Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.

Kann ich in dieser CUDA Academy-Lektion Code schreiben und ausführen?

Ja. Jede CUDA Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.

Alle Lektionen in diesem Kurs

  1. Warps, Lanes und Masken
  2. __shfl_down_sync für Reduktionen
  3. Ballot- und Vote-Funktionen
  4. Cooperative Groups
← Zurück zu CUDA Academy