السجلات والذاكرة المحلية
أسرع مساحة تخزين لكل خيط وتسريباتها إلى الذاكرة الأبطأ.
السجلات والذاكرة المحلية درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 1 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
The Fastest Memory You Have
Every thread gets its own private registers, the quickest storage on the chip. They live right next to the math units, so reads cost almost nothing.
Plain Variables Become Registers
When you write a normal local variable in a kernel, the compiler usually keeps it in a register. No special syntax is needed at all. 🙂
__global__ void k() {
int x = 5; // x lives in a register
float y = x * 2.0f;
}Private to One Thread
A register value belongs to exactly one thread. Thread 7 cannot see thread 8's register, which is why per-thread work stays cleanly separated.
Registers Are a Shared Pool
Each streaming multiprocessor has a fixed register file. All resident threads split it, so using fewer registers per thread lets more threads run at once.
When You Run Out
If a kernel needs more registers than are available, the extra values get pushed elsewhere. This event is called a register spill.
Where Spills Go: Local Memory
Spilled values land in local memory, which despite the name is not on-chip. It actually lives in slow off-chip DRAM.
Local Is Still Per-Thread
Local memory is private to each thread just like registers. The difference is purely speed, since local memory is far slower to reach.
Big Local Arrays Spill
Declaring a large per-thread array often forces it into local memory, because there simply are not enough registers to hold every element.
__global__ void k() {
float buf[64]; // likely lives in local memory
}Spills Hurt Performance
Every spill turns a free register access into a slow DRAM trip. Cutting register pressure is one of the easiest optimization wins you can make.
Seeing Your Register Count
You can ask the compiler to report registers per thread. Pass --ptxas-options=-v to nvcc and it prints usage for each kernel.
nvcc --ptxas-options=-v kernel.cuCapping Registers on Purpose
The __launch_bounds__ qualifier hints the compiler to limit registers, trading a little per-thread speed for more threads running together.
__global__ void __launch_bounds__(256)
myKernel() { /* ... */ }Quick Check
You declared a big per-thread array and performance dropped. What likely happened?
Recap: Fast and Private
You learned that registers are the fastest per-thread storage, that the pool is limited, and that overflow spills into slow local memory. Keep register pressure low. 🎯
الأسئلة الشائعة
هل درس «السجلات والذاكرة المحلية» مجاني؟
نعم — نص درس «السجلات والذاكرة المحلية» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.
ماذا ستتعلم في «السجلات والذاكرة المحلية»؟
أسرع مساحة تخزين لكل خيط وتسريباتها إلى الذاكرة الأبطأ. تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟
لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 1 من أصل 4.
كم من الوقت يستغرق درس «السجلات والذاكرة المحلية»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟
نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- السجلات والذاكرة المحلية
- مقايضات الذاكرة العامة
- الذاكرة الثابتة وذاكرة التخزين المؤقت الخاصة بها
- نموذج ذهني للهرمية