0Pricing
Mojo Academy · 课时

通过分块提升缓存局部性

将工作分块以适应缓存

通过分块提升缓存局部性 是 CoddyKit 上的免费 Mojo Academy 课时。 这是第 4 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Mojo Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Mojo Academy 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

The Cache Idea

The CPU keeps recently used data in a small, fast cache. Hits are quick; misses force a slow trip to main memory.

What Is Locality?

Cache locality means reusing data that is already nearby in the cache before it gets evicted to make room.

Big Data Spills the Cache

If your kernel sweeps a huge array, early elements are evicted before you reuse them, causing repeated cache misses.

Tile the Work

Tiling splits a big loop into small blocks that fit in cache, so each block's data stays hot while you use it.

alias tile = 64

An Outer and Inner Loop

Tiling turns one loop into two: an outer loop over blocks and an inner loop over the elements inside each block.

for t in range(0, n, tile):
    for i in range(t, min(t + tile, n)):
        out[i] = a[i] * 2

One Block at a Time

Each block is small enough to live in cache. You finish all its work before moving on, so reuse stays cheap.

Great for Matrices

Matrix multiply reuses rows and columns heavily. Blocking a matmul into tiles keeps reused data in cache and cuts misses.

Pick the Tile Size

The best tile fills the cache without overflowing it. Too small wastes reuse; too large spills, so you tune the size.

Tile Then Vectorize

Tiling and SIMD stack well. Vectorize the inner loop of each block to get cache locality and vector speed together.

for t in range(0, n, tile):
    vectorize[body, width](min(tile, n - t))

Mind the Edges

The last block may be smaller than a full tile. Clamp its range with min so you never read past the buffer.

var end = min(t + tile, n)

Measure the Gain

Tiling can help a lot or a little depending on sizes. Always benchmark a few tile values on your real data to choose.

Quick Check

Your kernel keeps missing cache because it sweeps a huge array end to end. What technique helps?

Recap

Tiling blocks a big loop so each chunk fits in cache; tune the tile size, vectorize the inner loop, and clamp the edges. 🧱

常见问题解答

「通过分块提升缓存局部性」课时是免费的吗?

是的 — 「通过分块提升缓存局部性」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Mojo Academy 课程的其余内容,请升级到 CoddyKit PRO。 Mojo Academy 课程共包含 4 节课。

「通过分块提升缓存局部性」这节课中我会学到什么?

将工作分块以适应缓存 你通过在浏览器中直接运行的动手代码来练习 Mojo Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Mojo Academy 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Mojo Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 4 节课,共 4 节。

「通过分块提升缓存局部性」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Mojo Academy 课中编写并运行代码吗?

能。每节 Mojo Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 计算内核的组成
  2. 结合 SIMD 与循环
  3. 减少内存流量
  4. 通过分块提升缓存局部性
← 返回 Mojo Academy