การสร้าง Matmul ทีละขั้นตอน
จากลูปพื้นฐานสู่เคอร์เนลจริง
การสร้าง Matmul ทีละขั้นตอน เป็นบทเรียน Mojo Academy ฟรีบน CoddyKit นี่คือบทเรียนที่ 2 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน Mojo Academy และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส Mojo Academy มีบทเรียนทั้งหมด 4 บทเรียน
บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ
What Matmul Computes
Matrix multiply, or matmul, combines an M by K matrix with a K by N matrix to produce an M by N result. It is the heart of AI math.
The Dot Product Rule
Each output cell is a dot product: walk across one row of A and down one column of B, multiplying and summing as you go.
Three Nested Loops
The naive version uses three loops over i, j, and k. The outer two pick the output cell, the inner one sums the products.
for i in range(M):
for j in range(N):
for k in range(K):
C[i, j] += A[i, k] * B[k, j]Start Each Cell at Zero
Before accumulating, set each result cell to zero. Otherwise old or garbage values pollute your sum.
C[i, j] = 0.0Accumulate the Sum
The inner k loop keeps a running accumulator. Summing into a local variable is often faster than touching C every step.
var acc: Float32 = 0.0
for k in range(K):
acc += A[i, k] * B[k, j]
C[i, j] = accUse fn for Speed
Write the kernel with fn and typed arguments. Strict types let Mojo compile tight machine code with no dynamic overhead.
fn matmul(A: Matrix, B: Matrix, C: Matrix):
passWhy Naive Is Slow
The basic triple loop does the right math but reads B by column, jumping through memory. Poor locality wastes cache and time.
Loop Order Matters
Reordering to i, k, j keeps the inner loop walking memory in straight lines. Better access patterns can speed matmul a lot.
for i in range(M):
for k in range(K):
for j in range(N):
C[i, j] += A[i, k] * B[k, j]The Inner Loop Is the Target
Nearly all the time lives in the innermost loop. That hot inner loop is exactly where vectorizing and tuning pay off.
Correctness First
Get the simple version right and save its output. It becomes the reference you compare every faster kernel against.
A Path to a Real Kernel
From here you add SIMD, tiling, and parallelism. Each step keeps the same result but raises throughput toward peak hardware speed.
Quick Check
Why is the textbook triple-loop matmul often slow in practice?
Recap
Matmul sums a dot product per output cell with three loops; start cells at zero, accumulate locally, and mind loop order for cache. 🔢
คำถามที่พบบ่อย
บทเรียน “การสร้าง Matmul ทีละขั้นตอน” ฟรีหรือไม่
ใช่ — ข้อความเต็มของ “การสร้าง Matmul ทีละขั้นตอน” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส Mojo Academy ให้อัปเกรดเป็น CoddyKit PRO คอร์ส Mojo Academy มีบทเรียนทั้งหมด 4 บทเรียน
คุณจะเรียนรู้อะไรในบทเรียน “การสร้าง Matmul ทีละขั้นตอน”
จากลูปพื้นฐานสู่เคอร์เนลจริง คุณปฏิบัติ Mojo Academy ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน
คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน Mojo Academy หรือไม่
ไม่จำเป็นต้องมีประสบการณ์มาก่อน Mojo Academy บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 2 จากทั้งหมด 4 บทเรียน
บทเรียน “การสร้าง Matmul ทีละขั้นตอน” ใช้เวลานานแค่ไหน
บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย
ฉันเขียนและรันโค้ดในบทเรียน Mojo Academy นี้ได้ไหม
ได้ บทเรียน Mojo Academy ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ
บทเรียนทั้งหมดในหลักสูตรนี้
- การจำลองเทนเซอร์ใน Mojo
- การสร้าง Matmul ทีละขั้นตอน
- การปรับปรุงผลคูณภายใน
- การตรวจสอบความถูกต้องเชิงตัวเลข