تداخل النسخ والحساب
أخفِ عمليات النقل خلف النوى.
تداخل النسخ والحساب درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
The Big Idea
Real speedups come from doing two things at once: copying one chunk of data while the GPU computes on another. 🔀
Separate Hardware Engines
A GPU has distinct copy engines and compute units. Overlap works because these can run truly in parallel, not just take turns.
Pinned Memory Is Required
Async copies need pinned host memory from cudaMallocHost. Ordinary pageable malloc forces a blocking copy that cannot overlap.
cudaMallocHost(&h_data, bytes);Async Copies
Use cudaMemcpyAsync with a stream so the transfer returns immediately and joins that stream's queue alongside kernels.
cudaMemcpyAsync(d_in, h_in, bytes, cudaMemcpyHostToDevice, s);Chunk the Work
Split the array into pieces and give each piece its own stream. While stream 0 computes, stream 1 can already be copying.
Copy, Compute, Copy Back
Each chunk does the same three steps in its stream: upload input, run the kernel, download output, all without blocking the host.
cudaMemcpyAsync(d, h, n, H2D, s);
k<<<g, b, 0, s>>>(d);
cudaMemcpyAsync(h, d, n, D2H, s);How the Overlap Forms
Because chunks live in different streams, the GPU can copy chunk two while still computing chunk one. The timeline fills up. 📊
Hiding the PCIe Cost
Transfers no longer add to total time; they hide behind compute. Ideally your runtime shrinks to roughly the longer of copy or kernel.
Watch the Default Stream
One stray default-stream call mid-loop can serialize everything again. Keep every async op tied to a real, non-default stream.
Verify in a Profiler
Open Nsight Systems and look for copy and compute lanes that overlap in time. Stacked, busy lanes mean your concurrency is real.
Synchronize at the End
After issuing all chunks, call cudaDeviceSynchronize once so every stream finishes before you read the final results on the host.
cudaDeviceSynchronize();Quick Check
What does an async transfer require to overlap with compute?
Recap
By chunking work across streams with pinned memory and async copies, you hide transfers behind compute and keep the GPU truly busy. 🚀
الأسئلة الشائعة
هل درس «تداخل النسخ والحساب» مجاني؟
نعم — نص درس «تداخل النسخ والحساب» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.
ماذا ستتعلم في «تداخل النسخ والحساب»؟
أخفِ عمليات النقل خلف النوى. تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟
لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.
كم من الوقت يستغرق درس «تداخل النسخ والحساب»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟
نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- مأزق Default Stream
- إنشاء Streams واستخدامها
- الأحداث لقياس الوقت والمزامنة
- تداخل النسخ والحساب