0Pricing
CUDA Academy · درس

مشكلة إعادة استخدام البيانات

لماذا تعيد النوى الساذجة قراءة الذاكرة العامة.

مشكلة إعادة استخدام البيانات درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 1 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

The Hidden Cost

A kernel can be correct yet slow because it keeps fetching the same data from slow global memory over and over. 🐢

Memory Is the Bottleneck

On the GPU, arithmetic is cheap but reaching global memory is expensive. Many kernels wait on memory far more than they compute.

Data Reuse Defined

Data reuse means one value loaded from global memory is used by many computations instead of being read again each time.

A Naive Stencil

Picture blurring an image. Each output pixel averages its neighbors, so every input pixel gets read by several different threads.

out[i] = (in[i-1] + in[i] + in[i+1]) / 3.0f;

Counting the Reads

In that blur, the value in[i] is read by threads i-1, i, and i+1. The same byte travels across the slow PCIe-fed bus three times.

Redundancy Adds Up

With a wider window or a 2D grid, each element may be re-read dozens of times. This redundant traffic dominates the runtime.

Bandwidth Is Finite

Global memory has a fixed peak bandwidth. Reading the same data repeatedly wastes that budget on bytes you already had.

Compute Sits Idle

While warps stall waiting on repeated global loads, the math units sit idle. You paid for cores you are barely using. 😴

Arithmetic Intensity

Arithmetic intensity is the ratio of math operations to bytes moved. Low intensity means memory, not compute, limits you.

The Goal: Read Once

The fix is to load each needed value once into fast on-chip storage, then let many threads reuse it from there.

Enter Shared Memory

That fast on-chip storage is shared memory. Staging data there is the foundation of every tiling optimization ahead.

Quick Check

Why is a naive stencil kernel often slow?

Recap

Naive kernels re-read shared data from global memory, wasting bandwidth and stalling compute. Tiling exists to load once and reuse. ✅

الأسئلة الشائعة

هل درس «مشكلة إعادة استخدام البيانات» مجاني؟

نعم — نص درس «مشكلة إعادة استخدام البيانات» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.

ماذا ستتعلم في «مشكلة إعادة استخدام البيانات»؟

لماذا تعيد النوى الساذجة قراءة الذاكرة العامة. تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟

لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 1 من أصل 4.

كم من الوقت يستغرق درس «مشكلة إعادة استخدام البيانات»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟

نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. مشكلة إعادة استخدام البيانات
  2. نمط التحميل والمزامنة والحساب
  3. القوالب ونافذات الانزلاق
  4. معالجة البلاطات الطرفية
← العودة إلى CUDA Academy