0Pricing
CUDA Academy · Lektion

Mit __syncthreads synchronisieren

Barrieren, die Threads im Gleichschritt halten.

Mit __syncthreads synchronisieren ist eine kostenlose CUDA Academy-Lektion auf CoddyKit. Dies ist Lektion 2 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des CUDA Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.

Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.

Threads Run Out of Step

Threads in a block do not march in lockstep. One may finish writing while another is still loading, so you cannot assume the data is ready yet.

Meet the Barrier

A barrier is a line in the code where every thread must wait until all of them arrive. Only then does the block move on together.

The __syncthreads Call

CUDA gives you __syncthreads() as the block-wide barrier. Call it and every thread in the block pauses until the last one shows up. 🛑

__syncthreads();

The Load-Sync-Use Order

The golden pattern is simple: every thread loads its piece into shared memory, you sync, then everyone safely reads what others wrote.

tile[threadIdx.x] = input[i];
__syncthreads();
float left = tile[threadIdx.x - 1];

Skip It and You Get Garbage

Forget the barrier and a thread may read a slot before its neighbor wrote it. That is a race condition, and the result is silent, wrong data.

It Is a Memory Fence Too

The barrier also makes shared writes visible to all threads after it. So once you pass __syncthreads, everyone sees the freshly written values.

All or None Must Reach It

The strict rule: every thread in the block must hit the same __syncthreads. If some skip it, the block can hang forever.

The Divergent Branch Trap

Never put a barrier inside an if that only some threads enter. Threads taking the other path never arrive, and the block deadlocks. ⚠️

if (threadIdx.x < 64) {
    __syncthreads();
}

Block Scope, Not Grid Scope

One important limit: __syncthreads only synchronizes one block. It cannot coordinate threads across different blocks of the grid.

Syncing Inside Loops

In tiled algorithms you often sync twice per phase: once after loading a tile and once after computing, before loading the next.

Cheap but Not Free

A barrier costs a little time while threads wait. Use it where correctness needs it, but avoid extra calls that just stall fast threads.

Quick Check

Let us test your grasp of the barrier.

Recap

You learned that __syncthreads() is a block-wide barrier that keeps threads in step, and that placing it inside divergent branches can deadlock. Next: bank conflicts. 🎯

Häufig gestellte Fragen

Ist die Lektion „Mit __syncthreads synchronisieren“ kostenlos?

Ja — der vollständige Text von „Mit __syncthreads synchronisieren“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des CUDA Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.

Was lerne ich in „Mit __syncthreads synchronisieren“?

Barrieren, die Threads im Gleichschritt halten. Du übst CUDA Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.

Brauche ich Erfahrung, um CUDA Academy zu starten?

Keine Vorkenntnisse erforderlich. CUDA Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 2 von 4.

Wie lange dauert die Lektion „Mit __syncthreads synchronisieren“?

Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.

Kann ich in dieser CUDA Academy-Lektion Code schreiben und ausführen?

Ja. Jede CUDA Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.

Alle Lektionen in diesem Kurs

  1. __shared__-Arrays deklarieren
  2. Mit __syncthreads synchronisieren
  3. Bankkonflikte vermeiden
  4. Dynamischer Shared Memory
← Zurück zu CUDA Academy