0Pricing
CUDA Academy · درس

ما هي معاملة الذاكرة

أسطر ذاكرة التخزين المؤقت والمقاطع ذات 32 أو 128 بايتًا.

ما هي معاملة الذاكرة درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 1 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

Memory Comes in Chunks

The GPU never fetches a single byte by itself. It moves memory in fixed-size blocks called a memory transaction, even if you only asked for one value.

The Warp Asks Together

A warp of 32 threads issues its loads at the same time. The hardware gathers all 32 requests and tries to serve them with as few transactions as possible.

Cache Lines and Segments

Global memory is carved into aligned segments, commonly 32 or 128 bytes wide. Each transaction pulls in one whole segment, never a partial slice.

Why 128 Bytes Matters

A 128-byte line holds exactly 32 four-byte floats. That is one value for every thread in a warp, so a tidy warp can be served by a single transaction.

// 32 threads x 4-byte float = 128 bytes = one line

Alignment Is Key

Segments start at fixed, aligned addresses. If your data begins mid-segment, a single warp can spill across two lines and force an extra transaction.

Paying for Bytes You Skip

If a warp touches only a few bytes of a 128-byte line, the GPU still moves all 128. The unused bytes are wasted bandwidth you paid for anyway.

Counting Transactions

Performance often comes down to one number: how many transactions a warp needs. Fewer transactions means you read less and finish faster.

The Best Case

When 32 threads read 32 neighboring floats from an aligned address, the GPU answers with one clean 128-byte line. That is the ideal single transaction.

float v = data[blockIdx.x * blockDim.x + threadIdx.x];

The Worst Case

If each thread grabs a value far from its neighbors, the warp may trigger up to 32 separate transactions, dragging in 32 full lines for tiny use.

Bandwidth Is the Limit

Most kernels are bound by how fast they move memory, not by math. Shrinking transaction count directly raises your effective bandwidth. 🚀

Set the Stage

Now you know the unit of cost: the segment-sized transaction. Every coalescing trick ahead is really about packing warps into as few of these as you can.

Quick Check

Test your grasp of transaction size.

Recap

You learned that the GPU moves memory in fixed segments, and a warp is served by as few transactions as it can manage. Fewer transactions, faster kernel. 🎉

الأسئلة الشائعة

هل درس «ما هي معاملة الذاكرة» مجاني؟

نعم — نص درس «ما هي معاملة الذاكرة» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.

ماذا ستتعلم في «ما هي معاملة الذاكرة»؟

أسطر ذاكرة التخزين المؤقت والمقاطع ذات 32 أو 128 بايتًا. تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟

لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 1 من أصل 4.

كم من الوقت يستغرق درس «ما هي معاملة الذاكرة»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟

نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. ما هي معاملة الذاكرة
  2. القراءات المتماسكة مقابل المتقطعة
  3. بنية المصفوفات مقابل مصفوفة البنى
  4. قياس عرض النطاق الفعلي
← العودة إلى CUDA Academy