0Pricing
CUDA Academy · درس

حلّل، حسّن، ثم أطلق

إغلاق الحلقة باستخدام مقاييس Nsight

حلّل، حسّن، ثم أطلق درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

Measure, Do Not Guess

Optimization starts with data, not hunches. Always profile first to learn where the time really goes before you change a single line. 🔍

Start With the Timeline

Open Nsight Systems to see the whole run: kernels, copies, and gaps. The timeline shows whether transfers and compute actually overlap.

Find the Hot Kernel

One stage usually dominates. The timeline reveals the longest-running kernel, and that is exactly where your tuning effort should go first.

Zoom In With Nsight Compute

For the hot kernel, switch to Nsight Compute. It reports per-kernel metrics like memory throughput, stalls, and achieved occupancy.

Read the Roofline

The roofline tells you if a kernel is memory-bound or compute-bound. That single answer decides which optimizations are even worth trying.

Fix Memory-Bound Kernels

If you are memory-bound, chase coalescing, shared-memory tiling, and vectorized loads. Feeding data faster is the only path to more speed.

Fix Compute-Bound Kernels

If you are compute-bound, raise ILP with unrolling, cut redundant math, or move to lower precision. Here the arithmetic units are the limit.

Label Your Code With NVTX

Wrap pipeline stages in NVTX ranges so the timeline shows named bars instead of anonymous kernels. Readable profiles are faster to debug.

nvtxRangePush("blur");
blurKernel<<<grid, block>>>(d_in, d_out);
nvtxRangePop();

Change One Thing at a Time

Apply a single optimization, then re-profile. This tight loop proves each change helped and stops you from chasing two effects at once.

Know When to Stop

When a kernel nears its roofline limit, more tuning yields little. Recognize diminishing returns and spend your time on the next bottleneck.

Validate, Then Ship

Before release, confirm correctness against a reference and check the speedup is stable. A fast but wrong pipeline is not shippable. 🚢

Quick Check

Nsight Compute says your hot kernel is memory-bound. Which fix should you try first?

Recap

You learned to profile with Nsight, read the roofline to classify kernels, apply targeted fixes one at a time, and validate before shipping a fast, correct pipeline. 🏁

الأسئلة الشائعة

هل درس «حلّل، حسّن، ثم أطلق» مجاني؟

نعم — نص درس «حلّل، حسّن، ثم أطلق» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.

ماذا ستتعلم في «حلّل، حسّن، ثم أطلق»؟

إغلاق الحلقة باستخدام مقاييس Nsight تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟

لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.

كم من الوقت يستغرق درس «حلّل، حسّن، ثم أطلق»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟

نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. تصميم مسار المعالجة
  2. دمج المرشحات في نواة واحدة
  3. بث البلاطات للصور الكبيرة
  4. حلّل، حسّن، ثم أطلق
← العودة إلى CUDA Academy