0Pricing
CUDA Academy · Lektion

Warp-Divergenz beseitigen

Indizieren Sie neu, damit die Warps ausgelastet bleiben.

Warp-Divergenz beseitigen ist eine kostenlose CUDA Academy-Lektion auf CoddyKit. Dies ist Lektion 2 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des CUDA Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.

Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.

Warps Run in Lockstep

A warp is 32 threads that execute the same instruction together. When their paths agree, the hardware runs at full speed.

What Divergence Costs

If threads in a warp take different branches, that is divergence. The hardware runs each path serially, leaving some lanes idle and wasting cycles.

The Naive Reduction Diverges

The simple version uses tid % (2*s) to pick active threads. Active and idle threads interleave inside every warp, so each warp diverges hard.

if (tid % (2 * s) == 0)
  data[tid] += data[tid + s];

Idle Lanes Still Cost

Even though half the threads do nothing, they still occupy the warp. The warp cannot finish until both the active and idle paths are handled.

Reindex by Thread ID

The fix is to map active work to the lowest thread IDs instead of scattered ones. Compute an index from tid and the stride.

int index = 2 * s * tid;
if (index < blockDim.x)
  data[index] += data[index + s];

Why That Helps

Now the busy threads are contiguous: tid 0,1,2,... all work, the rest all rest. Whole warps are either fully active or fully idle.

Fully Idle Warps Are Free

A warp where every lane is idle just retires with no work. There is no per-lane serialization, so the cost of divergence largely disappears.

The Modulo Trap

The hidden villain was the modulo condition. It scattered active threads across each warp, which is exactly what creates divergence.

Same Work, Better Mapping

You did not change the math or the number of additions. You only remapped which thread does each add, and the warps thank you for it.

It Compounds at Scale

Across thousands of blocks and many steps, removing divergence is a real speedup, often a couple of times faster than the naive kernel.

Still One Snag Left

This version reads neighbors that are interleaved in shared memory, which can cause bank conflicts. The next lesson fixes that too.

Quick Check

Think about what causes warp divergence in the naive reduction.

Recap

You killed divergence by giving work to the lowest thread IDs, so warps are all-active or all-idle. Same math, faster reduction. Next: bank conflicts. 🚀

Häufig gestellte Fragen

Ist die Lektion „Warp-Divergenz beseitigen“ kostenlos?

Ja — der vollständige Text von „Warp-Divergenz beseitigen“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des CUDA Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.

Was lerne ich in „Warp-Divergenz beseitigen“?

Indizieren Sie neu, damit die Warps ausgelastet bleiben. Du übst CUDA Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.

Brauche ich Erfahrung, um CUDA Academy zu starten?

Keine Vorkenntnisse erforderlich. CUDA Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 2 von 4.

Wie lange dauert die Lektion „Warp-Divergenz beseitigen“?

Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.

Kann ich in dieser CUDA Academy-Lektion Code schreiben und ausführen?

Ja. Jede CUDA Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.

Alle Lektionen in diesem Kurs

  1. Die Idee des Reduktionsbaums
  2. Warp-Divergenz beseitigen
  3. Sequenzielle Adressierung
  4. Abschließende Reduktion über mehrere Blöcke
← Zurück zu CUDA Academy