计算内核的组成
执行实际工作的热点内层循环
计算内核的组成 是 CoddyKit 上的免费 Mojo Academy 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Mojo Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Mojo Academy 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
What Is a Kernel?
A compute kernel is the small, focused routine that does the real numerical work, like adding two arrays element by element.
The Hot Inner Loop
Most of a kernel's time lives in one tight inner loop. Speed up that loop and you speed up the whole program.
for i in range(n):
out[i] = a[i] + b[i]Inputs and Outputs
A clean kernel takes its data buffers as parameters, so the same routine can run on any arrays you pass in.
fn add(a: UnsafePointer[Float32], b: UnsafePointer[Float32], out: UnsafePointer[Float32], n: Int):
passUse fn for Strictness
Kernels are written with fn, not def. The strict, typed style lets Mojo compile tight machine code with no surprises.
fn kernel(n: Int):
passKeep the Body Small
The fastest kernels do one thing. A short, predictable inner body is easy for the compiler to optimize aggressively.
No Surprises Inside
Avoid heavy work like allocation or I/O in the hot loop. Each iteration should be cheap and uniform for peak throughput.
Count the Work
Think in terms of total operations. A kernel over n elements does roughly n units of work, so n drives the cost.
Memory Is the Limit
Many kernels are not compute-bound but memory-bound. They wait on data, so how you move bytes often matters most.
Plan to Vectorize
Design the loop so each step can later process many elements with SIMD. A regular access pattern makes vectorizing easy.
A Tiny Saxpy Kernel
A classic kernel is saxpy: out = scale times a plus b. It is small, regular, and a great baseline to optimize.
for i in range(n):
out[i] = scale * a[i] + b[i]Measure Before You Tune
Start with the simple correct version and time it. That number is your baseline for every optimization that follows.
Quick Check
You are about to optimize a kernel. Where does almost all of its time go?
Recap
A kernel is a small fn whose tight inner loop does the work; keep it uniform, mind memory, and measure a baseline first. 🔧
常见问题解答
「计算内核的组成」课时是免费的吗?
是的 — 「计算内核的组成」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Mojo Academy 课程的其余内容,请升级到 CoddyKit PRO。 Mojo Academy 课程共包含 4 节课。
「计算内核的组成」这节课中我会学到什么?
执行实际工作的热点内层循环 你通过在浏览器中直接运行的动手代码来练习 Mojo Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Mojo Academy 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Mojo Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。
「计算内核的组成」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Mojo Academy 课中编写并运行代码吗?
能。每节 Mojo Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- 计算内核的组成
- 结合 SIMD 与循环
- 减少内存流量
- 通过分块提升缓存局部性