0Pricing
CUDA Academy · درس

استخدام printf داخل النواة

اعرض المخرجات من خيوط الجهاز.

استخدام printf داخل النواة درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 3 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

Printing from the GPU

Believe it or not, you can call printf right inside a kernel. It is the simplest way to peek at what your threads are doing. 👀

__global__ void hi() {
    printf("Hello from the GPU\n");
}

Every Thread Prints

Remember the kernel runs in every thread, so a single printf line fires once per thread. Launch 256 threads and you get 256 lines.

hi<<<1, 256>>>(); // 256 hellos

Identify Each Thread

Include the threadIdx in your message so you can tell threads apart. Otherwise the output is just a wall of identical lines.

printf("thread %d\n", threadIdx.x);

Output Order Is Not Fixed

Threads run in parallel, so the printed order is unpredictable. Do not rely on lines arriving in index sequence.

Format Strings Work

Device printf supports the usual format specifiers like %d, %f, and %s. It feels just like host printf.

printf("i=%d val=%f\n", i, x[i]);

Output Goes to a Buffer

Device output is staged in a GPU buffer and flushed to your console later, not the instant printf runs. That is normal.

You Must Wait to See It

Because the launch is async, you will not see prints until the GPU finishes. Call cudaDeviceSynchronize to flush them.

hi<<<1, 4>>>();
cudaDeviceSynchronize();

Guard Heavy Printing

Printing from millions of threads floods the buffer. Guard it so only thread 0 prints, or only a few do.

if (threadIdx.x == 0) printf("block done\n");

Great for Quick Debugging

printf is your fastest debugging tool: drop one in to check an index or a value, confirm the bug, then remove it.

printf("i=%d should be < n=%d\n", i, n);

It Slows Kernels Down

Heavy printing hurts performance badly. Use it to find a problem, then delete it before you measure real speed.

Old GPUs May Differ

Device printf needs a reasonably modern compute capability (2.0 and up). Almost every current GPU supports it just fine.

Quick Check

Check what you know about device printf.

Recap: Kernel printf

Use printf to peek inside threads, add threadIdx to tell them apart, sync to flush, and remove it before timing. Handy tool! 🎉

الأسئلة الشائعة

هل درس «استخدام printf داخل النواة» مجاني؟

نعم — نص درس «استخدام printf داخل النواة» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.

ماذا ستتعلم في «استخدام printf داخل النواة»؟

اعرض المخرجات من خيوط الجهاز. تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟

لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 3 من أصل 4.

كم من الوقت يستغرق درس «استخدام printf داخل النواة»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟

نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. تشريح النواة
  2. التشغيل بأقواس الزوايا الثلاثية
  3. استخدام printf داخل النواة
  4. شرح cudaDeviceSynchronize
← العودة إلى CUDA Academy