0Pricing
CUDA Academy · درس

واجهة API لأجزاء WMMA

تحميل الأجزاء وتشغيل mma_sync وتخزينها

واجهة API لأجزاء WMMA درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 3 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

The WMMA API

To program Tensor Cores directly you use WMMA, the warp matrix multiply-accumulate API in the nvcuda::wmma namespace. 🧩

A Whole Warp Cooperates

WMMA is warp-wide: all 32 threads in a warp work together on one tile. You think in tiles, not in single threads.

Meet the Fragment

A fragment is the WMMA data type holding one warp's slice of a matrix tile. Each thread quietly owns a few of its elements.

Three Fragment Roles

You declare fragments tagged matrix_a, matrix_b, or accumulator. The tag tells WMMA how that tile feeds the multiply-accumulate.

Tile Shapes Are Fixed

Fragments use set tile sizes like 16 by 16 by 16. You pick a supported shape; the hardware only accepts those combinations.

Step 1: Load

First you call load_matrix_sync to pull a tile from memory into a fragment. This load spreads the data across the warp for you.

wmma::load_matrix_sync(a_frag, ptr, ldm);

Step 2: Clear the Accumulator

Before adding, you zero the result tile with fill_fragment. A clean accumulator means your sums start from a known value.

wmma::fill_fragment(c_frag, 0.0f);

Step 3: Multiply-Accumulate

The core call is mma_sync. It runs D = A times B plus C on the loaded fragments using the Tensor Cores directly.

wmma::mma_sync(c_frag, a_frag, b_frag, c_frag);

Step 4: Store

Finally store_matrix_sync writes the accumulator fragment back to memory. This store gathers each thread's pieces into one tile.

wmma::store_matrix_sync(out, c_frag, ldm, layout);

Sync Means Warp-Synchronous

The _sync suffix is a reminder: every WMMA call is warp-synchronous. All 32 lanes must reach it together or behavior is undefined.

Layout and Leading Dimension

Loads and stores need the leading dimension, the stride between rows, plus a row- or column-major layout to read memory correctly.

Quick Check

Which WMMA call actually runs the matrix multiply-accumulate on the Tensor Cores?

Recap

You walked the WMMA flow: declare a fragment, load tiles, clear the accumulator, call mma_sync, then store. Every step is warp-synchronous. 🙌

الأسئلة الشائعة

هل درس «واجهة API لأجزاء WMMA» مجاني؟

نعم — نص درس «واجهة API لأجزاء WMMA» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.

ماذا ستتعلم في «واجهة API لأجزاء WMMA»؟

تحميل الأجزاء وتشغيل mma_sync وتخزينها تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟

لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 3 من أصل 4.

كم من الوقت يستغرق درس «واجهة API لأجزاء WMMA»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟

نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. ما الذي تحسبه Tensor Cores
  2. الدقة المختلطة: FP16 وBF16 وTF32
  3. واجهة API لأجزاء WMMA
  4. مفاضلات الاستقرار العددي
← العودة إلى CUDA Academy