نموذج ذهني للهرمية
طابق البيانات مع مساحة الذاكرة المناسبة.
نموذج ذهني للهرمية درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
A Pyramid of Tradeoffs
Think of GPU memory as a pyramid. The top is tiny and lightning fast, while the base is huge but slow. Your job is to use each layer well.
Top: Registers
At the peak sit registers: per-thread, fastest, and very limited. Keep your hottest working values here whenever you possibly can.
Next: Shared Memory
Just below comes shared memory, an on-chip scratchpad visible to all threads in a block. It is your tool for fast intra-block teamwork.
Side Path: Constant Cache
Alongside sits the constant cache, ideal for small read-only values that a warp reads uniformly and broadcasts in one fetch.
Base: Global Memory
The wide base is global memory: gigabytes, visible to everyone, but high-latency. It holds your big inputs and outputs.
The Trap: Local Memory
Watch out for local memory. It sounds close by but lives in slow DRAM, used only when registers spill. Avoid relying on it.
Scope Decides the Space
Pick by who needs the data. One thread alone wants registers, a block working together wants shared memory, everyone wants global.
Lifetime Decides Too
Registers and shared memory vanish when a kernel ends, but global memory persists across launches. Match storage to how long data must live.
The Golden Rule
The winning pattern is load once from global, compute in fast on-chip memory, then write once back. This minimizes slow global memory traffic.
Capacity Costs Occupancy
Spending lots of registers or shared memory per block lets fewer blocks run at once. Occupancy is the balance you constantly tune.
Putting It Together
Great kernels deliberately route each piece of data to the right layer. That single habit is what separates slow code from fast CUDA code. 💪
Quick Check
Two threads in the SAME block need to share intermediate results. Which space fits best?
Recap: Match Data to Layer
You built a mental map: registers, shared, constant, and global trade speed for size and scope. Routing data to the right layer is the whole game. 🧠
الأسئلة الشائعة
هل درس «نموذج ذهني للهرمية» مجاني؟
نعم — نص درس «نموذج ذهني للهرمية» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.
ماذا ستتعلم في «نموذج ذهني للهرمية»؟
طابق البيانات مع مساحة الذاكرة المناسبة. تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟
لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.
كم من الوقت يستغرق درس «نموذج ذهني للهرمية»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟
نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- السجلات والذاكرة المحلية
- مقايضات الذاكرة العامة
- الذاكرة الثابتة وذاكرة التخزين المؤقت الخاصة بها
- نموذج ذهني للهرمية