การผสาน SIMD กับลูป
ทำแกนหลักของเคอร์เนลให้เป็นเวกเตอร์
การผสาน SIMD กับลูป เป็นบทเรียน Mojo Academy ฟรีบน CoddyKit นี่คือบทเรียนที่ 2 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน Mojo Academy และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส Mojo Academy มีบทเรียนทั้งหมด 4 บทเรียน
บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ
From Scalar to SIMD
To accelerate a kernel, swap the one-at-a-time loop for one that handles a whole pack of values per step with SIMD.
for i in range(n):
out[i] = a[i] + b[i]Pick a Width
Choose how many elements fit in one vector. The width is the number of lanes each step processes together.
alias width = 4Load a Chunk
Grab several elements at once into a SIMD value with a vector load instead of reading them one by one.
var va = a.load[width=width](i)Compute the Pack
Run the kernel's math on both chunks at once. The whole pack is added element-wise in a single operation.
var vsum = a.load[width=width](i) + b.load[width=width](i)Store the Pack
Write the full result back with a vector store, covering every lane you just computed in one move.
out.store[width=width](i, vsum)Step by the Width
The vectorized loop advances by the width, not by one. Each pass covers a full pack of elements.
for i in range(0, n, width):
passLet vectorize Help
Mojo's vectorize helper sweeps a closure across the range in vector steps so you skip the manual bookkeeping.
from algorithm import vectorizeWrite the Chunk Closure
You define a small parameterized function that handles one chunk. Mojo calls it with the right width as it sweeps.
fn body[w: Int](i: Int):
out.store[width=w](i, a.load[width=w](i) + b.load[width=w](i))Run vectorize
Call vectorize with your closure, the width, and the size. It loops and even handles the leftover tail for you.
vectorize[body, width](n)Mind the Tail
When n is not a multiple of the width, a few elements remain. vectorize cleans up that tail so nothing is missed.
Same Output, More Speed
The vectorized kernel produces identical results but moves through data in big steps, so it finishes much sooner.
Quick Check
You vectorize a kernel with width 4 but n is 10. What handles the last two elements?
Recap
Vectorize a kernel by loading and storing packs, stepping by the width, and letting vectorize handle the sweep and the tail. 🚀
คำถามที่พบบ่อย
บทเรียน “การผสาน SIMD กับลูป” ฟรีหรือไม่
ใช่ — ข้อความเต็มของ “การผสาน SIMD กับลูป” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส Mojo Academy ให้อัปเกรดเป็น CoddyKit PRO คอร์ส Mojo Academy มีบทเรียนทั้งหมด 4 บทเรียน
คุณจะเรียนรู้อะไรในบทเรียน “การผสาน SIMD กับลูป”
ทำแกนหลักของเคอร์เนลให้เป็นเวกเตอร์ คุณปฏิบัติ Mojo Academy ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน
คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน Mojo Academy หรือไม่
ไม่จำเป็นต้องมีประสบการณ์มาก่อน Mojo Academy บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 2 จากทั้งหมด 4 บทเรียน
บทเรียน “การผสาน SIMD กับลูป” ใช้เวลานานแค่ไหน
บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย
ฉันเขียนและรันโค้ดในบทเรียน Mojo Academy นี้ได้ไหม
ได้ บทเรียน Mojo Academy ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ
บทเรียนทั้งหมดในหลักสูตรนี้
- กายวิภาคของเคอร์เนลคำนวณ
- การผสาน SIMD กับลูป
- การลดการรับส่งข้อมูลในหน่วยความจำ
- การแบ่งบล็อกเพื่อความใกล้เคียงของแคช