หลีกเลี่ยงความขัดแย้งของแบงก์
ทำความเข้าใจว่าการเติมช่องว่างช่วยให้เข้าถึงหน่วยความจำร่วมเร็วขึ้นได้อย่างไร
หลีกเลี่ยงความขัดแย้งของแบงก์ เป็นบทเรียน CUDA Academy ฟรีบน CoddyKit นี่คือบทเรียนที่ 3 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน CUDA Academy และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส CUDA Academy มีบทเรียนทั้งหมด 4 บทเรียน
บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ
Shared Memory Has Banks
Shared memory is split into 32 banks, one for each thread in a warp. Spreading accesses across them lets all 32 threads read at full speed.
How Addresses Map
Consecutive 4-byte words land in consecutive banks, wrapping around after 32. So word 0 is bank 0, word 32 is bank 0 again.
What a Bank Conflict Is
A bank conflict happens when two threads in a warp hit the same bank but different words. The hardware must serialize those accesses.
Serialization Costs Speed
An N-way conflict turns one fast access into N slow ones. A 32-way conflict can make shared memory as slow as if it were single-file. 🐢
The Happy Case: Stride One
When each thread reads tile[threadIdx.x], every thread hits a distinct bank. That is conflict-free and runs at full bandwidth.
float v = tile[threadIdx.x];The Trap: Stride of 32
Reading with a stride of 32 sends every thread to the same bank. That is the worst case, a full 32-way conflict.
float v = tile[threadIdx.x * 32];Even Strides Hurt Too
Any even stride that shares a factor with 32 causes partial conflicts. A stride of 2, for example, gives a 2-way conflict across the warp.
The Broadcast Exception
Good news: if all threads read the same address, the hardware broadcasts it in one shot. Same word is fine, only same bank with different words conflicts.
Padding to the Rescue
For 2D tiles, add one extra column with +1 padding. This shifts each row so column accesses no longer all land in one bank.
__shared__ float tile[32][33];Why +1 Works
The extra column makes the row length coprime with 32. Now stepping down a column visits a different bank each time, killing the conflict.
Profile, Do Not Guess
Bank conflicts are invisible in source code. Use Nsight Compute to measure shared-memory conflicts before spending effort fixing them.
Quick Check
Let us check what makes shared access fast.
Recap
You learned that shared memory has 32 banks, that same-bank different-word access serializes, and that +1 padding fixes column conflicts. Next: dynamic shared memory. 🎯
คำถามที่พบบ่อย
บทเรียน “หลีกเลี่ยงความขัดแย้งของแบงก์” ฟรีหรือไม่
ใช่ — ข้อความเต็มของ “หลีกเลี่ยงความขัดแย้งของแบงก์” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส CUDA Academy ให้อัปเกรดเป็น CoddyKit PRO คอร์ส CUDA Academy มีบทเรียนทั้งหมด 4 บทเรียน
คุณจะเรียนรู้อะไรในบทเรียน “หลีกเลี่ยงความขัดแย้งของแบงก์”
ทำความเข้าใจว่าการเติมช่องว่างช่วยให้เข้าถึงหน่วยความจำร่วมเร็วขึ้นได้อย่างไร คุณปฏิบัติ CUDA Academy ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน
คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน CUDA Academy หรือไม่
ไม่จำเป็นต้องมีประสบการณ์มาก่อน CUDA Academy บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 3 จากทั้งหมด 4 บทเรียน
บทเรียน “หลีกเลี่ยงความขัดแย้งของแบงก์” ใช้เวลานานแค่ไหน
บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย
ฉันเขียนและรันโค้ดในบทเรียน CUDA Academy นี้ได้ไหม
ได้ บทเรียน CUDA Academy ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ
บทเรียนทั้งหมดในหลักสูตรนี้
- ประกาศอาร์เรย์ __shared__
- ซิงโครไนซ์ด้วย __syncthreads
- หลีกเลี่ยงความขัดแย้งของแบงก์
- หน่วยความจำร่วมแบบไดนามิก