0Pricing
CUDA Academy · Lektion

Die WMMA-Fragment-API

Fragmente laden, mit mma_sync berechnen und speichern

Die WMMA-Fragment-API ist eine kostenlose CUDA Academy-Lektion auf CoddyKit. Dies ist Lektion 3 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des CUDA Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.

Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.

The WMMA API

To program Tensor Cores directly you use WMMA, the warp matrix multiply-accumulate API in the nvcuda::wmma namespace. 🧩

A Whole Warp Cooperates

WMMA is warp-wide: all 32 threads in a warp work together on one tile. You think in tiles, not in single threads.

Meet the Fragment

A fragment is the WMMA data type holding one warp's slice of a matrix tile. Each thread quietly owns a few of its elements.

Three Fragment Roles

You declare fragments tagged matrix_a, matrix_b, or accumulator. The tag tells WMMA how that tile feeds the multiply-accumulate.

Tile Shapes Are Fixed

Fragments use set tile sizes like 16 by 16 by 16. You pick a supported shape; the hardware only accepts those combinations.

Step 1: Load

First you call load_matrix_sync to pull a tile from memory into a fragment. This load spreads the data across the warp for you.

wmma::load_matrix_sync(a_frag, ptr, ldm);

Step 2: Clear the Accumulator

Before adding, you zero the result tile with fill_fragment. A clean accumulator means your sums start from a known value.

wmma::fill_fragment(c_frag, 0.0f);

Step 3: Multiply-Accumulate

The core call is mma_sync. It runs D = A times B plus C on the loaded fragments using the Tensor Cores directly.

wmma::mma_sync(c_frag, a_frag, b_frag, c_frag);

Step 4: Store

Finally store_matrix_sync writes the accumulator fragment back to memory. This store gathers each thread's pieces into one tile.

wmma::store_matrix_sync(out, c_frag, ldm, layout);

Sync Means Warp-Synchronous

The _sync suffix is a reminder: every WMMA call is warp-synchronous. All 32 lanes must reach it together or behavior is undefined.

Layout and Leading Dimension

Loads and stores need the leading dimension, the stride between rows, plus a row- or column-major layout to read memory correctly.

Quick Check

Which WMMA call actually runs the matrix multiply-accumulate on the Tensor Cores?

Recap

You walked the WMMA flow: declare a fragment, load tiles, clear the accumulator, call mma_sync, then store. Every step is warp-synchronous. 🙌

Häufig gestellte Fragen

Ist die Lektion „Die WMMA-Fragment-API“ kostenlos?

Ja — der vollständige Text von „Die WMMA-Fragment-API“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des CUDA Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.

Was lerne ich in „Die WMMA-Fragment-API“?

Fragmente laden, mit mma_sync berechnen und speichern Du übst CUDA Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.

Brauche ich Erfahrung, um CUDA Academy zu starten?

Keine Vorkenntnisse erforderlich. CUDA Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 3 von 4.

Wie lange dauert die Lektion „Die WMMA-Fragment-API“?

Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.

Kann ich in dieser CUDA Academy-Lektion Code schreiben und ausführen?

Ja. Jede CUDA Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.

Alle Lektionen in diesem Kurs

  1. Was Tensor Cores berechnen
  2. Gemischte Genauigkeit: FP16, BF16, TF32
  3. Die WMMA-Fragment-API
  4. Abwägungen bei der numerischen Stabilität
← Zurück zu CUDA Academy