การเข้าถึงหน่วยความจำแบบเพียร์ทูเพียร์
คัดลอกข้อมูลโดยตรงระหว่าง GPU ผ่าน NVLink
การเข้าถึงหน่วยความจำแบบเพียร์ทูเพียร์ เป็นบทเรียน CUDA Academy ฟรีบน CoddyKit นี่คือบทเรียนที่ 3 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน CUDA Academy และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส CUDA Academy มีบทเรียนทั้งหมด 4 บทเรียน
บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ
The Slow Detour
Moving data from one GPU to another usually bounces through the CPU's memory first. That round trip is slow and wastes the host's bandwidth.
GPUs Talking Directly
Modern GPUs can skip the CPU entirely. Peer-to-peer access lets one GPU read and write another's memory over a direct link.
The NVLink Highway
The fast link between cards is often NVLink, far quicker than the shared PCIe bus. P2P transfers ride this highway when it is available.
Check Before You Trust
Not every pair of GPUs can do P2P. Ask first with cudaDeviceCanAccessPeer, which reports whether one device may reach another.
int can;
cudaDeviceCanAccessPeer(&can, 0, 1);Turn On Access
Permission is off by default. While GPU 0 is current, call cudaDeviceEnablePeerAccess to let it reach into GPU 1's memory.
cudaSetDevice(0);
cudaDeviceEnablePeerAccess(1, 0);Access Is One-Way
Enabling peer access grants only the direction you ask for. For GPU 1 to read GPU 0, you must enable that direction too, from device 1.
Copying Peer to Peer
To move a buffer directly between cards, use cudaMemcpyPeer. You name the destination pointer and device plus the source pointer and device.
cudaMemcpyPeer(dst, 1, src, 0, bytes);Kernels Reading Across
With access enabled, a kernel on GPU 0 can dereference a pointer that lives on GPU 1. The hardware fetches it over the peer link transparently.
When P2P Is Not There
If two cards cannot peer, the runtime quietly falls back to staging through host memory. You still get a copy, just at the slower PCIe rate.
Async Peer Copies
Peer transfers can overlap other work. Pair cudaMemcpyPeerAsync with a stream so the copy runs while kernels keep computing.
cudaMemcpyPeerAsync(dst, 1, src, 0, bytes, stream);Turn It Off When Done
Peer access uses resources, so release it when finished. cudaDeviceDisablePeerAccess tears down the link you opened earlier.
cudaDeviceDisablePeerAccess(1);Quick Check
Recall the benefit of peer-to-peer access between two GPUs.
Recap
P2P lets GPUs share memory directly, ideally over NVLink. Check support, enable access per direction, then use cudaMemcpyPeer. Next: scaling with NCCL. ✨
คำถามที่พบบ่อย
บทเรียน “การเข้าถึงหน่วยความจำแบบเพียร์ทูเพียร์” ฟรีหรือไม่
ใช่ — ข้อความเต็มของ “การเข้าถึงหน่วยความจำแบบเพียร์ทูเพียร์” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส CUDA Academy ให้อัปเกรดเป็น CoddyKit PRO คอร์ส CUDA Academy มีบทเรียนทั้งหมด 4 บทเรียน
คุณจะเรียนรู้อะไรในบทเรียน “การเข้าถึงหน่วยความจำแบบเพียร์ทูเพียร์”
คัดลอกข้อมูลโดยตรงระหว่าง GPU ผ่าน NVLink คุณปฏิบัติ CUDA Academy ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน
คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน CUDA Academy หรือไม่
ไม่จำเป็นต้องมีประสบการณ์มาก่อน CUDA Academy บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 3 จากทั้งหมด 4 บทเรียน
บทเรียน “การเข้าถึงหน่วยความจำแบบเพียร์ทูเพียร์” ใช้เวลานานแค่ไหน
บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย
ฉันเขียนและรันโค้ดในบทเรียน CUDA Academy นี้ได้ไหม
ได้ บทเรียน CUDA Academy ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ
บทเรียนทั้งหมดในหลักสูตรนี้
- การแจกแจงและเลือกอุปกรณ์
- การแบ่งงานข้าม GPU
- การเข้าถึงหน่วยความจำแบบเพียร์ทูเพียร์
- หลาย GPU ด้วย NCCL