قياس التسارع
قارن أداء التنفيذ الساذج بالتنفيذ المقسّم إلى بلاطات.
قياس التسارع درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
Prove the Win
You built a tiled kernel, but how much faster is it really? Measuring turns a guess into a number you can trust. 📊
Time on the GPU's Clock
Use CUDA events to time kernels. They sit in the GPU stream and measure exactly when work starts and finishes.
cudaEvent_t start, stop;
cudaEventCreate(&start);
cudaEventCreate(&stop);Bracket the Kernel
Record an event before the launch and another after, then synchronize so the timing actually waits for the GPU to finish.
cudaEventRecord(start);
matmul<<<g,b>>>(...);
cudaEventRecord(stop);
cudaEventSynchronize(stop);Read the Elapsed Time
Ask for the gap between events in milliseconds. That single number is your kernel's measured runtime.
float ms;
cudaEventElapsedTime(&ms, start, stop);Warm Up First
The very first launch pays one-time setup costs. Run a warm-up kernel before timing so you measure steady-state speed, not startup.
Average Several Runs
One sample is noisy. Time the kernel several times and take the average for a stable, honest result.
Compute Speedup
Speedup is simply naive time divided by tiled time. A value of 4x means the tiled kernel ran four times faster.
float speedup = naive_ms / tiled_ms;Think in GFLOPS
Matmul does about 2 * N^3 floating-point operations. Divide that by your time to report performance in GFLOPS, the standard yardstick.
double gflops = (2.0*N*N*N) / (ms * 1e6);Why Tiling Wins
The speedup comes from slashing global memory traffic. Shared-memory reuse keeps the cores fed instead of waiting on slow loads.
Always Verify Correctness
A fast wrong answer is useless. Compare your GPU result against a CPU reference before you celebrate the speedup. ✅
Know Your Ceiling
Even tiled matmul trails hand-tuned cuBLAS. Knowing the gap tells you when to optimize further and when to call a library.
Quick Check
Recall the right tool for timing GPU kernels.
Recap
You timed with CUDA events, warmed up, averaged, computed speedup and GFLOPS, and verified correctness. You now prove your optimizations. 🏁
الأسئلة الشائعة
هل درس «قياس التسارع» مجاني؟
نعم — نص درس «قياس التسارع» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.
ماذا ستتعلم في «قياس التسارع»؟
قارن أداء التنفيذ الساذج بالتنفيذ المقسّم إلى بلاطات. تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟
لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.
كم من الوقت يستغرق درس «قياس التسارع»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟
نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.