0Pricing
Mojo Academy · บทเรียน

การผสานการทำงานแบบขนานกับเวกเตอร์

เพิ่มการทำงานแบบเธรดบน SIMD

การผสานการทำงานแบบขนานกับเวกเตอร์ เป็นบทเรียน Mojo Academy ฟรีบน CoddyKit นี่คือบทเรียนที่ 3 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน Mojo Academy และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส Mojo Academy มีบทเรียนทั้งหมด 4 บทเรียน

บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ

Two Kinds of Speed

SIMD packs many values into one core's instruction. Threads spread work across cores. Stacking both gives you the most. ⚡

Outer Layer: Threads

Use parallelize on the outer level so each core owns a chunk of the data. That is the coarse split across cores.

parallelize[do_chunk](workers)

Inner Layer: Vectors

Inside each chunk, use vectorize so one core chews through its slice in SIMD packs. That is the fine split per core.

from algorithm import vectorize

Import Both Helpers

Both tools live in the algorithm module, so you can bring them in together and combine them in one function.

from algorithm import parallelize, vectorize

The Chunk Worker

Each chunk worker computes its start and end, then hands its own slice to vectorize for the SIMD sweep.

fn do_chunk(c: Int):
    var start = c * chunk

Vectorize the Inner Range

Within the worker, call vectorize on just this chunk's length. The SIMD width sets how many lanes run at once.

vectorize[step, width](end - start)

Offset Into the Data

The inner step gets a local index, so add the chunk start to reach the right spot in the full array.

fn step[w: Int](i: Int):
    out.store[width=w](start + i, ...)

Pick the SIMD Width

Match the vector width to your hardware lanes with simdwidthof so each core uses its registers fully.

alias width = simdwidthof[DType.float32]()

Cores Times Lanes

The win multiplies: many cores, each doing many lanes per step. That product is why combined code is so fast.

Keep Chunks Independent

This stacking only works because chunks never touch each other's data. Independence keeps the two layers safe.

Measure the Combined Gain

Benchmark scalar, vector-only, and combined versions. The numbers reveal how much each layer contributes.

Quick Check

You combine threads and SIMD on the same workload.

Recap

You wrap parallelize over chunks and call vectorize inside each, offsetting by the chunk start, so cores times lanes multiply your throughput. 🚀

คำถามที่พบบ่อย

บทเรียน “การผสานการทำงานแบบขนานกับเวกเตอร์” ฟรีหรือไม่

ใช่ — ข้อความเต็มของ “การผสานการทำงานแบบขนานกับเวกเตอร์” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส Mojo Academy ให้อัปเกรดเป็น CoddyKit PRO คอร์ส Mojo Academy มีบทเรียนทั้งหมด 4 บทเรียน

คุณจะเรียนรู้อะไรในบทเรียน “การผสานการทำงานแบบขนานกับเวกเตอร์”

เพิ่มการทำงานแบบเธรดบน SIMD คุณปฏิบัติ Mojo Academy ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน

คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน Mojo Academy หรือไม่

ไม่จำเป็นต้องมีประสบการณ์มาก่อน Mojo Academy บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 3 จากทั้งหมด 4 บทเรียน

บทเรียน “การผสานการทำงานแบบขนานกับเวกเตอร์” ใช้เวลานานแค่ไหน

บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย

ฉันเขียนและรันโค้ดในบทเรียน Mojo Academy นี้ได้ไหม

ได้ บทเรียน Mojo Academy ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ

บทเรียนทั้งหมดในหลักสูตรนี้

  1. ฟังก์ชัน parallelize
  2. การแบ่งงานเป็นชิ้น
  3. การผสานการทำงานแบบขนานกับเวกเตอร์
  4. การหลีกเลี่ยงภาวะแข่งขันข้อมูล
← กลับไปที่ Mojo Academy