0Pricing
CUDA Academy · درس

دورة حياة برنامج CUDA

احجز الذاكرة، وانسخ البيانات، وشغّل النواة، وانسخ النتائج، ثم حرّر الذاكرة.

دورة حياة برنامج CUDA درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

A Repeating Rhythm

Almost every CUDA program follows the same five-beat dance. Learn the rhythm once and you can read any GPU code. 🕺

Step 1: Allocate

First reserve device memory for your inputs and outputs with cudaMalloc, since the GPU cannot use host buffers directly.

cudaMalloc(&d_a, bytes);
cudaMalloc(&d_b, bytes);

Step 2: Copy In

Next upload your input data from host to device with cudaMemcpy. The GPU now has its own copy to work on.

cudaMemcpy(d_a, h_a, bytes, cudaMemcpyHostToDevice);

Step 3: Launch

Now launch your kernel across many threads. This is where the parallel work actually happens on the device. 🚀

myKernel<<<blocks, threads>>>(d_a, d_b, n);

Step 4: Copy Back

When compute finishes, download the results from device to host so the CPU can read and use them.

cudaMemcpy(h_b, d_b, bytes, cudaMemcpyDeviceToHost);

Step 5: Free

Finally release every device buffer with cudaFree. Skipping this leaks GPU memory that other work could use.

cudaFree(d_a);
cudaFree(d_b);

Don't Forget to Sync

Because launches are async, call cudaDeviceSynchronize before reading results to ensure the GPU has truly finished.

cudaDeviceSynchronize();

Copies Cost Time

The copy steps cross the slow PCIe bus, so data transfer is often the real bottleneck, not the kernel itself.

Reuse Buffers

Allocating once and reusing buffers across many launches beats allocating fresh memory every iteration of a loop.

Symmetry of the Pattern

Notice the symmetry: every cudaMalloc pairs with a cudaFree, and every copy in eventually pairs with a copy out.

The Lifecycle in One Breath

Say it like a mantra: allocate, copy, launch, copy back, free. That single line is the skeleton of nearly every kernel program.

Quick Check

Let us check the standard CUDA program lifecycle.

Recap

You learned the five-step CUDA lifecycle: allocate, copy in, launch, copy back, free, plus syncing before you read results. 🎉

الأسئلة الشائعة

هل درس «دورة حياة برنامج CUDA» مجاني؟

نعم — نص درس «دورة حياة برنامج CUDA» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.

ماذا ستتعلم في «دورة حياة برنامج CUDA»؟

احجز الذاكرة، وانسخ البيانات، وشغّل النواة، وانسخ النتائج، ثم حرّر الذاكرة. تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟

لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.

كم من الوقت يستغرق درس «دورة حياة برنامج CUDA»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟

نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. مؤهّل الدالة __global__
  2. دوال __device__ و__host__
  3. مساحات عناوين منفصلة
  4. دورة حياة برنامج CUDA
← العودة إلى CUDA Academy