0Pricing
CUDA Academy · درس

الوصول إلى الذاكرة من نظير إلى نظير

نسخ مباشرة بين وحدات GPU عبر NVLink

الوصول إلى الذاكرة من نظير إلى نظير درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 3 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

The Slow Detour

Moving data from one GPU to another usually bounces through the CPU's memory first. That round trip is slow and wastes the host's bandwidth.

GPUs Talking Directly

Modern GPUs can skip the CPU entirely. Peer-to-peer access lets one GPU read and write another's memory over a direct link.

The NVLink Highway

The fast link between cards is often NVLink, far quicker than the shared PCIe bus. P2P transfers ride this highway when it is available.

Check Before You Trust

Not every pair of GPUs can do P2P. Ask first with cudaDeviceCanAccessPeer, which reports whether one device may reach another.

int can;
cudaDeviceCanAccessPeer(&can, 0, 1);

Turn On Access

Permission is off by default. While GPU 0 is current, call cudaDeviceEnablePeerAccess to let it reach into GPU 1's memory.

cudaSetDevice(0);
cudaDeviceEnablePeerAccess(1, 0);

Access Is One-Way

Enabling peer access grants only the direction you ask for. For GPU 1 to read GPU 0, you must enable that direction too, from device 1.

Copying Peer to Peer

To move a buffer directly between cards, use cudaMemcpyPeer. You name the destination pointer and device plus the source pointer and device.

cudaMemcpyPeer(dst, 1, src, 0, bytes);

Kernels Reading Across

With access enabled, a kernel on GPU 0 can dereference a pointer that lives on GPU 1. The hardware fetches it over the peer link transparently.

When P2P Is Not There

If two cards cannot peer, the runtime quietly falls back to staging through host memory. You still get a copy, just at the slower PCIe rate.

Async Peer Copies

Peer transfers can overlap other work. Pair cudaMemcpyPeerAsync with a stream so the copy runs while kernels keep computing.

cudaMemcpyPeerAsync(dst, 1, src, 0, bytes, stream);

Turn It Off When Done

Peer access uses resources, so release it when finished. cudaDeviceDisablePeerAccess tears down the link you opened earlier.

cudaDeviceDisablePeerAccess(1);

Quick Check

Recall the benefit of peer-to-peer access between two GPUs.

Recap

P2P lets GPUs share memory directly, ideally over NVLink. Check support, enable access per direction, then use cudaMemcpyPeer. Next: scaling with NCCL. ✨

الأسئلة الشائعة

هل درس «الوصول إلى الذاكرة من نظير إلى نظير» مجاني؟

نعم — نص درس «الوصول إلى الذاكرة من نظير إلى نظير» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.

ماذا ستتعلم في «الوصول إلى الذاكرة من نظير إلى نظير»؟

نسخ مباشرة بين وحدات GPU عبر NVLink تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟

لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 3 من أصل 4.

كم من الوقت يستغرق درس «الوصول إلى الذاكرة من نظير إلى نظير»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟

نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. حصر الأجهزة واختيارها
  2. تقسيم العمل بين وحدات GPU
  3. الوصول إلى الذاكرة من نظير إلى نظير
  4. استخدام وحدات GPU متعددة مع NCCL
← العودة إلى CUDA Academy