Transformer 架构:注意力、词元与上下文
您将追踪自注意力机制,理解 BERT 如何一次性读取完整句子而不是从左到右读取,并解读 CLS 和 SEP 这两个特殊词元。
Transformer 架构:注意力、词元与上下文 是 CoddyKit 上的免费 Machine Learning Academy 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Machine Learning Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Machine Learning Academy 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
What Is a Transformer?
The Transformer is a neural network architecture introduced in 2017 that replaced recurrent networks for most NLP tasks. Unlike RNNs that process tokens one at a time, Transformers process the entire sequence in parallel using a mechanism called self-attention. This parallel processing makes training much faster and allows the model to capture long-range dependencies more effectively.
Self-Attention: Relating Every Token
Self-attention allows each token in a sequence to attend to every other token simultaneously. For the sentence 'The bank by the river was steep', the word 'bank' can attend strongly to 'river' to resolve its meaning. Each token produces three vectors: Query (Q), Key (K), and Value (V), which are used to compute weighted relationships between all token pairs.
import torch
import torch.nn.functional as F
# Simplified self-attention for 3 tokens, d_model=4
Q = torch.randn(3, 4) # queries
K = torch.randn(3, 4) # keys
V = torch.randn(3, 4) # values
d_k = Q.shape[-1]
scores = torch.matmul(Q, K.T) / (d_k ** 0.5) # scaled dot product
weights = F.softmax(scores, dim=-1) # attention weights
output = torch.matmul(weights, V) # weighted values
print('Attention weights:', weights)Scaled Dot-Product Attention
The attention score between token i and token j is computed as the dot product of Q_i and K_j, divided by the square root of the key dimension to prevent vanishingly small gradients. The formula is: Attention(Q, K, V) = softmax(QK^T / sqrt(d_k)) * V. The scaling factor sqrt(d_k) keeps gradients stable for large embedding dimensions.
Multi-Head Attention
Instead of one set of Q, K, V projections, Transformers use multi-head attention: h parallel attention heads, each learning different aspects of token relationships. One head might learn syntactic dependencies (subject-verb), another semantic ones (synonyms). The outputs of all heads are concatenated and projected to produce the final representation.
import torch.nn as nn
multihead_attn = nn.MultiheadAttention(
embed_dim=512,
num_heads=8, # 8 heads, each with dim 64
dropout=0.1,
batch_first=True
)
# x shape: (batch, seq_len, 512)
# output shape: (batch, seq_len, 512)
output, attn_weights = multihead_attn(x, x, x)BERT: Bidirectional Context
BERT (Bidirectional Encoder Representations from Transformers) reads the entire sequence at once, attending to both left and right context simultaneously. Earlier models like GPT read left-to-right only. This bidirectionality lets BERT understand that 'bank' in 'river bank' differs from 'bank' in 'bank account' by seeing all surrounding words at once.
Special Tokens: CLS and SEP
BERT introduces two special tokens. The [CLS] (classification) token is prepended to every input; after processing, its final hidden state aggregates sentence-level information and is used for classification tasks. The [SEP] token separates two sentences in tasks like question answering or next-sentence prediction. Understanding these tokens is essential when building BERT pipelines.
# Example tokenised input for BERT sentence-pair
# [CLS] I love Python [SEP] Python is great [SEP]
# token_ids: [101, 1045, 2293, 18750, 102, 18750, 2003, 2307, 102]
# segment_ids: [0, 0, 0, 0, 0, 1, 1, 1, 1 ]
print('CLS token id:', 101)
print('SEP token id:', 102)Positional Encoding: Order Without Recurrence
Because Transformers process all tokens in parallel, they have no inherent sense of token order. Positional encodings are added to each token embedding to inject position information. BERT uses learned positional embeddings while the original Transformer used sinusoidal functions. Without positional encoding, 'cat bites dog' and 'dog bites cat' would produce identical representations.
import torch.nn as nn
# BERT-style learned positional embedding
pos_embedding = nn.Embedding(512, 768) # max 512 positions, d_model=768
positions = torch.arange(seq_len).unsqueeze(0) # (1, seq_len)
pos_enc = pos_embedding(positions) # (1, seq_len, 768)
# Added to token embeddings before feeding to transformer layersEncoder Architecture: Layers and Feed-Forward
Each BERT encoder layer consists of two sub-layers: multi-head self-attention followed by a position-wise feed-forward network (two linear layers with a GELU activation). Each sub-layer has a residual connection and layer normalisation. BERT-base stacks 12 such layers; BERT-large uses 24. Deeper stacks capture more abstract linguistic structure.
Token Embeddings: WordPiece Vocabulary
BERT tokenises text using WordPiece subword tokenisation. Rare words are split into frequent sub-units: 'unbelievable' might become ['un', '##believe', '##able']. The ## prefix indicates a continuation subword. This approach handles out-of-vocabulary words gracefully and uses a vocabulary of ~30,000 tokens, balancing coverage and embedding table size.
from transformers import BertTokenizer
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
text = 'unbelievable achievements'
tokens = tokenizer.tokenize(text)
print(tokens) # ['un', '##believ', '##able', 'achievements']
encoded = tokenizer(text, return_tensors='pt')
print('input_ids:', encoded['input_ids'])
print('attention_mask:', encoded['attention_mask'])Attention Mask: Handling Padding
When processing batches of variable-length sentences, shorter sentences are padded with [PAD] tokens to match the longest sequence. The attention mask is a binary tensor (1 for real tokens, 0 for padding) that tells the model to ignore padding positions in the attention computation. Without this mask, the model would attend to meaningless padding tokens and corrupt its representations.
from transformers import BertTokenizer
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
batch = ['Short text.', 'This sentence is longer than the first one.']
encoded = tokenizer(batch, padding=True, truncation=True, return_tensors='pt')
print('input_ids shape:', encoded['input_ids'].shape)
print('attention_mask:\n', encoded['attention_mask'])
# Zeros mark padding positionsPre-Training BERT: MLM and NSP
BERT was pre-trained on two tasks. Masked Language Modelling (MLM) randomly masks 15% of tokens and trains BERT to predict the original token from context, forcing bidirectional understanding. Next Sentence Prediction (NSP) trains BERT to determine whether two sentences are consecutive, helping sentence-pair tasks. Fine-tuning then adapts these rich representations to downstream tasks with minimal additional training.
Quick Check
Test your understanding of Machine Learning with Python concepts from this lesson.
Lesson Recap
In this lesson you learned: Transformers use parallel self-attention instead of sequential recurrence, BERT reads bidirectional context using masked language modelling pre-training, and special tokens [CLS] and [SEP] structure BERT's inputs for classification and sentence-pair tasks. Next up we explore how Hugging Face tokenizers encode raw text into the tensor format BERT expects.
常见问题解答
「Transformer 架构:注意力、词元与上下文」课时是免费的吗?
是的 — 「Transformer 架构:注意力、词元与上下文」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Machine Learning Academy 课程的其余内容,请升级到 CoddyKit PRO。 Machine Learning Academy 课程共包含 4 节课。
「Transformer 架构:注意力、词元与上下文」这节课中我会学到什么?
您将追踪自注意力机制,理解 BERT 如何一次性读取完整句子而不是从左到右读取,并解读 CLS 和 SEP 这两个特殊词元。 你通过在浏览器中直接运行的动手代码来练习 Machine Learning Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Machine Learning Academy 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Machine Learning Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。
「Transformer 架构:注意力、词元与上下文」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Machine Learning Academy 课中编写并运行代码吗?
能。每节 Machine Learning Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- Transformer 架构:注意力、词元与上下文
- Hugging Face 分词器:为 BERT 编码文本
- 微调 BertForSequenceClassification
- 评估与推理:从 Logits 到预测标签