0Pricing
CUDA Academy · บทเรียน

ข้อแลกเปลี่ยนของหน่วยความจำส่วนกลาง

มีขนาดใหญ่ ช้า และทุกเธรดเข้าถึงได้

ข้อแลกเปลี่ยนของหน่วยความจำส่วนกลาง เป็นบทเรียน CUDA Academy ฟรีบน CoddyKit นี่คือบทเรียนที่ 2 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน CUDA Academy และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส CUDA Academy มีบทเรียนทั้งหมด 4 บทเรียน

บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ

The Biggest Space You Have

Global memory is the GPU's main DRAM, often many gigabytes. It is where your big input and output arrays naturally live during a kernel.

Visible to Everyone

Every thread in every block can read and write global memory. That shared visibility is exactly why you copy your data there before launching. 🌍

You Allocate It with cudaMalloc

You reserve global memory from the host with cudaMalloc, which hands you a device pointer to that DRAM region.

float *d_a;
cudaMalloc(&d_a, n * sizeof(float));

Large but Slow

The catch is latency. A single global memory read can cost hundreds of clock cycles, far more than a register access ever would.

Bandwidth Is the Real Limit

Many kernels are bound not by math but by how fast bytes flow from DRAM. We call these memory-bound kernels.

Hiding Latency with Threads

The GPU hides slow reads by switching to other ready warps while one waits. This latency hiding is why having many threads matters so much.

It Persists Across Launches

Data in global memory stays put between kernel launches until you free it. You can run several kernels over the same buffers.

Cached by L2

A unified L2 cache sits in front of global memory for the whole GPU. Reused addresses can be served from it instead of slow DRAM.

Access Pattern Decides Speed

How threads map to addresses hugely affects throughput. Neighboring threads touching neighboring addresses is what later lessons call coalescing.

Minimize the Trips

The core strategy is simple: touch global memory as few times as possible. Read once, reuse on-chip, then write once.

Free What You Allocate

Because global memory is a finite resource, release it with cudaFree when you are done to avoid leaking device memory.

cudaFree(d_a);

Quick Check

Which statement best captures the tradeoff of global memory?

Recap: Big, Shared, Slow

You now know global memory is the GPU's large, all-visible DRAM that trades capacity for high latency. Touch it rarely and reuse data on-chip. 🚀

คำถามที่พบบ่อย

บทเรียน “ข้อแลกเปลี่ยนของหน่วยความจำส่วนกลาง” ฟรีหรือไม่

ใช่ — ข้อความเต็มของ “ข้อแลกเปลี่ยนของหน่วยความจำส่วนกลาง” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส CUDA Academy ให้อัปเกรดเป็น CoddyKit PRO คอร์ส CUDA Academy มีบทเรียนทั้งหมด 4 บทเรียน

คุณจะเรียนรู้อะไรในบทเรียน “ข้อแลกเปลี่ยนของหน่วยความจำส่วนกลาง”

มีขนาดใหญ่ ช้า และทุกเธรดเข้าถึงได้ คุณปฏิบัติ CUDA Academy ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน

คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน CUDA Academy หรือไม่

ไม่จำเป็นต้องมีประสบการณ์มาก่อน CUDA Academy บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 2 จากทั้งหมด 4 บทเรียน

บทเรียน “ข้อแลกเปลี่ยนของหน่วยความจำส่วนกลาง” ใช้เวลานานแค่ไหน

บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย

ฉันเขียนและรันโค้ดในบทเรียน CUDA Academy นี้ได้ไหม

ได้ บทเรียน CUDA Academy ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ

บทเรียนทั้งหมดในหลักสูตรนี้

  1. รีจิสเตอร์และหน่วยความจำเฉพาะที่
  2. ข้อแลกเปลี่ยนของหน่วยความจำส่วนกลาง
  3. หน่วยความจำคงที่และแคชของมัน
  4. แบบจำลองทางความคิดของลำดับชั้น
← กลับไปที่ CUDA Academy