การผสานการทำงานแบบขนานกับเวกเตอร์
เพิ่มการทำงานแบบเธรดบน SIMD
การผสานการทำงานแบบขนานกับเวกเตอร์ เป็นบทเรียน Mojo Academy ฟรีบน CoddyKit นี่คือบทเรียนที่ 3 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน Mojo Academy และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส Mojo Academy มีบทเรียนทั้งหมด 4 บทเรียน
บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ
Two Kinds of Speed
SIMD packs many values into one core's instruction. Threads spread work across cores. Stacking both gives you the most. ⚡
Outer Layer: Threads
Use parallelize on the outer level so each core owns a chunk of the data. That is the coarse split across cores.
parallelize[do_chunk](workers)Inner Layer: Vectors
Inside each chunk, use vectorize so one core chews through its slice in SIMD packs. That is the fine split per core.
from algorithm import vectorizeImport Both Helpers
Both tools live in the algorithm module, so you can bring them in together and combine them in one function.
from algorithm import parallelize, vectorizeThe Chunk Worker
Each chunk worker computes its start and end, then hands its own slice to vectorize for the SIMD sweep.
fn do_chunk(c: Int):
var start = c * chunkVectorize the Inner Range
Within the worker, call vectorize on just this chunk's length. The SIMD width sets how many lanes run at once.
vectorize[step, width](end - start)Offset Into the Data
The inner step gets a local index, so add the chunk start to reach the right spot in the full array.
fn step[w: Int](i: Int):
out.store[width=w](start + i, ...)Pick the SIMD Width
Match the vector width to your hardware lanes with simdwidthof so each core uses its registers fully.
alias width = simdwidthof[DType.float32]()Cores Times Lanes
The win multiplies: many cores, each doing many lanes per step. That product is why combined code is so fast.
Keep Chunks Independent
This stacking only works because chunks never touch each other's data. Independence keeps the two layers safe.
Measure the Combined Gain
Benchmark scalar, vector-only, and combined versions. The numbers reveal how much each layer contributes.
Quick Check
You combine threads and SIMD on the same workload.
Recap
You wrap parallelize over chunks and call vectorize inside each, offsetting by the chunk start, so cores times lanes multiply your throughput. 🚀
คำถามที่พบบ่อย
บทเรียน “การผสานการทำงานแบบขนานกับเวกเตอร์” ฟรีหรือไม่
ใช่ — ข้อความเต็มของ “การผสานการทำงานแบบขนานกับเวกเตอร์” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส Mojo Academy ให้อัปเกรดเป็น CoddyKit PRO คอร์ส Mojo Academy มีบทเรียนทั้งหมด 4 บทเรียน
คุณจะเรียนรู้อะไรในบทเรียน “การผสานการทำงานแบบขนานกับเวกเตอร์”
เพิ่มการทำงานแบบเธรดบน SIMD คุณปฏิบัติ Mojo Academy ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน
คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน Mojo Academy หรือไม่
ไม่จำเป็นต้องมีประสบการณ์มาก่อน Mojo Academy บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 3 จากทั้งหมด 4 บทเรียน
บทเรียน “การผสานการทำงานแบบขนานกับเวกเตอร์” ใช้เวลานานแค่ไหน
บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย
ฉันเขียนและรันโค้ดในบทเรียน Mojo Academy นี้ได้ไหม
ได้ บทเรียน Mojo Academy ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ
บทเรียนทั้งหมดในหลักสูตรนี้
- ฟังก์ชัน parallelize
- การแบ่งงานเป็นชิ้น
- การผสานการทำงานแบบขนานกับเวกเตอร์
- การหลีกเลี่ยงภาวะแข่งขันข้อมูล