缩放点积与多头注意力
在并行的子空间中进行注意
缩放点积与多头注意力 是 CoddyKit 上的免费 Deep Learning Academy 课时。 这是第 2 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Deep Learning Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Deep Learning Academy 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
The Scaling Problem
When key vectors are long, dot-product scores grow huge and softmax becomes razor-sharp. That kills gradients, so we need to scale the scores down.
Divide by Root d
The fix is to divide each score by the square root of the key dimension. This keeps the numbers in a stable range before softmax.
scores = (q @ k.transpose(-2, -1)) / d_k ** 0.5Scaled Dot-Product Attention
Put it together: score, scale, softmax, then blend values. This four-step recipe is the famous scaled dot-product attention.
w = scores.softmax(dim=-1)
out = w @ vOne Head, One View
A single attention head learns just one way to relate tokens. But language has many relationships at once, so one head is limiting.
Many Heads in Parallel
Multi-head attention runs several attention heads side by side, each with its own Q, K, V projections, each attending in a different subspace.
Split the Dimension
You do not make the model wider. You split the model dimension across heads, so each head works on a smaller slice in parallel.
d_head = d_model // num_headsReshape into Heads
To run heads in parallel, you reshape Q, K, V to add a heads axis, then attend within each head independently.
q = q.view(B, T, H, d_head).transpose(1, 2)Different Patterns
One head may track grammar, another may link a pronoun to its noun. Multiple heads capture diverse patterns the same input contains.
Concatenate Heads
After each head produces its output, you concatenate them back into one vector of the original model dimension.
out = out.transpose(1, 2).reshape(B, T, d_model)Final Projection
A last output projection mixes the concatenated head results, letting the model combine what every head discovered.
out = self.W_o(out)Use the Built-in
In practice PyTorch ships nn.MultiheadAttention, so you rarely hand-roll the reshapes. Knowing the math still helps you debug.
attn = nn.MultiheadAttention(d_model, num_heads)Quick Check
Let's confirm why we scale the scores.
Recap
You saw how scaling stabilizes softmax and how multi-head attention splits the work across parallel heads, then concatenates and projects. Great progress!
常见问题解答
「缩放点积与多头注意力」课时是免费的吗?
是的 — 「缩放点积与多头注意力」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Deep Learning Academy 课程的其余内容,请升级到 CoddyKit PRO。 Deep Learning Academy 课程共包含 4 节课。
「缩放点积与多头注意力」这节课中我会学到什么?
在并行的子空间中进行注意 你通过在浏览器中直接运行的动手代码来练习 Deep Learning Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Deep Learning Academy 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Deep Learning Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 2 节课,共 4 节。
「缩放点积与多头注意力」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Deep Learning Academy 课中编写并运行代码吗?
能。每节 Deep Learning Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。