مؤشر واحد على الجانبين
تعرّف على كيفية عمل cudaMallocManaged.
مؤشر واحد على الجانبين درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 1 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
The Two-Pointer Headache
So far you juggled two pointers: one for the host, one for the device. Unified Memory replaces that pain with a single pointer both sides can use. 🙂
Meet cudaMallocManaged
You allocate managed memory with cudaMallocManaged. It hands back one address that works in your CPU code and inside your kernels alike.
float *data;
cudaMallocManaged(&data, n * sizeof(float));No More cudaMemcpy
The big win: you usually skip the manual copies. With managed memory the runtime moves data for you, so cudaMemcpy often disappears from your code.
Write on the Host
You can fill the buffer with ordinary CPU code right after allocating. The same pointer you got is a normal address your loops can touch.
for (int i = 0; i < n; i++)
data[i] = i;Read on the Device
Pass that exact pointer to your kernel and the GPU threads dereference it directly. One address serves both worlds with no translation step.
kernel<<<blocks, threads>>>(data, n);Sync Before You Read Back
After a kernel writes managed data, call cudaDeviceSynchronize before the CPU reads it. That guarantees the GPU finished and results are visible.
kernel<<<b, t>>>(data, n);
cudaDeviceSynchronize();Free It Like Any Buffer
Managed memory is still device memory, so you release it with cudaFree. There is no special managed-free call to remember.
cudaFree(data);Less Boilerplate, Fewer Bugs
Because you delete the alloc-copy-launch-copy-free dance, your programs shrink. Fewer copies means fewer chances to mix up directions or sizes.
Great for Prototyping
Unified Memory is perfect when you want a kernel running fast. You prototype quickly, then optimize transfers later only where they actually matter.
It Is Not Free Magic
The data still has to travel across PCIe under the hood. Convenience is real, but performance can lag hand-tuned copies until you add hints later.
When to Reach for It
Choose managed memory for simpler code, deep pointer structures, or oversubscribing GPU memory. It shines when clarity matters more than raw peak speed.
Quick Check
Let us confirm how managed allocation differs from the classic flow.
Recap: One Pointer, Both Sides
You learned that cudaMallocManaged hands you one pointer for host and device, dropping most copies. Sync before reading, free with cudaFree. Nice work! 🎉
الأسئلة الشائعة
هل درس «مؤشر واحد على الجانبين» مجاني؟
نعم — نص درس «مؤشر واحد على الجانبين» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.
ماذا ستتعلم في «مؤشر واحد على الجانبين»؟
تعرّف على كيفية عمل cudaMallocManaged. تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟
لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 1 من أصل 4.
كم من الوقت يستغرق درس «مؤشر واحد على الجانبين»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟
نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- مؤشر واحد على الجانبين
- ترحيل الصفحات عند الطلب
- الجلب المسبق باستخدام cudaMemPrefetchAsync
- التلميحات عبر cudaMemAdvise