Autograd:用于反向传播的自动微分
您将定义一个标量计算图,调用 .backward(),并检查叶张量上的 .grad,以理解梯度如何在网络中流动。
Autograd:用于反向传播的自动微分 是 CoddyKit 上的免费 Machine Learning Academy 课时。 这是第 2 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Machine Learning Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Machine Learning Academy 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
What Is Automatic Differentiation?
Automatic differentiation (autograd) is the engine that computes gradients in PyTorch without requiring the programmer to derive them by hand. Unlike numerical differentiation (finite differences) or symbolic differentiation (algebra), autograd works by recording operations on tensors at runtime and replaying them in reverse. This makes it possible to train models of any architecture efficiently and correctly.
import torch
# A simple differentiable computation
x = torch.tensor(3.0, requires_grad=True)
y = x ** 2 + 2 * x + 1 # y = (x+1)^2
# Compute gradient dy/dx
y.backward()
print(x.grad) # tensor(8.) because dy/dx = 2x+2 = 2*3+2 = 8requires_grad: Enabling Gradient Tracking
Gradient tracking is disabled by default. Setting requires_grad=True tells PyTorch to record all operations on that tensor so gradients can be computed later. Leaf tensors (parameters) require gradients; intermediate tensors inherit gradient-tracking from their parents automatically. Model parameters created by nn.Module have requires_grad=True set automatically.
import torch
# Leaf tensor with gradient tracking
w = torch.tensor(2.0, requires_grad=True)
b = torch.tensor(0.5, requires_grad=True)
# Forward pass
x = torch.tensor(3.0) # input, no grad needed
y_pred = w * x + b # y = wx + b = 6.5
print(y_pred.requires_grad) # True (inherited)
print(w.is_leaf) # True
print(y_pred.is_leaf) # FalseThe Computation Graph
As you perform operations on tensors with requires_grad=True, PyTorch builds a dynamic computation graph (a directed acyclic graph of operations). Each node stores the operation and references to its inputs. When .backward() is called, PyTorch traverses this graph in reverse (from output to inputs) using the chain rule to compute partial derivatives for every leaf tensor. The graph is re-built fresh on each forward pass, enabling dynamic architectures.
import torch
a = torch.tensor(2.0, requires_grad=True)
b = torch.tensor(3.0, requires_grad=True)
# Build computation graph
c = a * b # c = 6
d = c + a # d = 8
e = d ** 2 # e = 64
# Inspect graph node
print(e.grad_fn) # <PowBackward0 object>
print(e.grad_fn.next_functions) # parents in the graphCalling .backward(): Computing Gradients
Calling .backward() on a scalar tensor triggers backpropagation through the entire computation graph, depositing gradients into the .grad attribute of each leaf tensor. If the output is not scalar, you must provide a gradient tensor of the same shape (the upstream gradient). Gradients accumulate by default — always call optimizer.zero_grad() (or tensor.grad.zero_()) before the next backward pass to avoid incorrect updates.
import torch
x = torch.tensor(4.0, requires_grad=True)
loss = (x - 1) ** 2 # minimum at x=1
loss.backward()
print(x.grad) # tensor(6.) = 2*(x-1) = 2*3 = 6
# Gradients accumulate! Reset before next pass
x.grad.zero_()
new_loss = (x - 2) ** 2
new_loss.backward()
print(x.grad) # tensor(4.) = 2*(x-2) = 2*2 = 4Gradient Flow Through Multiple Operations
The chain rule states that the derivative of a composition of functions is the product of derivatives. Autograd applies this mechanically through every operation in the graph. Understanding how gradients flow helps debug issues like vanishing gradients (products near zero) and exploding gradients (products greater than 1 repeated many times). Activation function choice (ReLU vs sigmoid) directly affects gradient magnitudes.
import torch
# Chain: z = sigmoid(wx + b), loss = (z - y_true)^2
def sigmoid(x):
return 1 / (1 + torch.exp(-x))
w = torch.tensor(0.5, requires_grad=True)
b = torch.tensor(0.0, requires_grad=True)
x = torch.tensor(2.0)
y_true = torch.tensor(1.0)
z = sigmoid(w * x + b)
loss = (z - y_true) ** 2
loss.backward()
print(f'dL/dw = {w.grad:.4f}') # gradient w.r.t. weight
print(f'dL/db = {b.grad:.4f}') # gradient w.r.t. biastorch.no_grad(): Disabling Gradient Tracking
During inference (prediction on new data), you do not need gradients — computing them wastes time and memory. Wrapping inference code with torch.no_grad() disables the computation graph entirely, speeding up forward passes and reducing memory usage by up to 50%. This is also used during evaluation loops to get validation metrics without influencing training.
import torch
x = torch.randn(100, requires_grad=True)
# With gradient tracking (training)
out = x ** 2
print(out.requires_grad) # True
# Without gradient tracking (inference)
with torch.no_grad():
out_no_grad = x ** 2
print(out_no_grad.requires_grad) # False
# Also used as a decorator
@torch.no_grad()
def predict(model, data):
return model(data)Gradient Accumulation Issue and zero_grad
PyTorch accumulates (adds) gradients into .grad each time .backward() is called. This is intentional for some advanced techniques, but during normal training it means you must zero gradients before each backward pass. The standard idiom uses the optimizer's zero_grad() method, which clears gradients for all parameters the optimizer manages. Forgetting this step leads to gradients growing unboundedly, corrupting updates.
import torch
import torch.nn as nn
import torch.optim as optim
model = nn.Linear(2, 1)
optimizer = optim.SGD(model.parameters(), lr=0.01)
for epoch in range(3):
optimizer.zero_grad() # clear old gradients
x = torch.randn(4, 2)
y_pred = model(x)
loss = y_pred.sum()
loss.backward() # compute new gradients
optimizer.step() # update weights
print(f'Epoch {epoch}: loss={loss.item():.3f}')Inspecting Gradient Values for Debugging
Inspecting gradient values is crucial for diagnosing training problems. You can access them via tensor.grad after calling backward. Plotting the gradient norm across layers reveals whether gradients vanish (near zero) or explode (very large) in deep networks. A common debugging hook is to print gradient statistics after every N batches to detect instability early in training.
import torch
import torch.nn as nn
model = nn.Sequential(
nn.Linear(4, 8),
nn.ReLU(),
nn.Linear(8, 1)
)
x = torch.randn(16, 4)
y_pred = model(x)
loss = y_pred.mean()
loss.backward()
# Inspect gradients of each parameter
for name, param in model.named_parameters():
if param.grad is not None:
norm = param.grad.norm().item()
print(f'{name}: grad_norm={norm:.4f}')Detaching Tensors from the Graph
Sometimes you want to use a computed tensor as a constant input to another computation — without letting gradients flow through it. .detach() returns a new tensor that shares data but is removed from the computation graph. This is used in target networks (reinforcement learning), stopping gradient flow in specific branches of a network, and converting tensors to NumPy arrays for plotting.
import torch
a = torch.tensor(3.0, requires_grad=True)
b = a * 2 # b depends on a
c = b.detach() # c is a constant copy of b's value
d = c * 5 # gradient does NOT flow back to a
d.backward()
# a.grad is None because c severed the graph
print(a.grad) # None
# Detach before converting to NumPy
np_val = b.detach().numpy()
print(np_val) # [6.0] (as NumPy array)Second-Order Gradients with create_graph
PyTorch supports higher-order differentiation. Passing create_graph=True to .backward() keeps the computation graph for the gradient computation itself, allowing you to differentiate through gradients. This is used in meta-learning (MAML), gradient penalty regularisation (Wasserstein GAN), and any technique that optimises with respect to gradients.
import torch
x = torch.tensor(2.0, requires_grad=True)
y = x ** 3 # y = x^3
# First derivative dy/dx = 3x^2
dy = torch.autograd.grad(y, x, create_graph=True)[0]
print(dy) # tensor(12., grad_fn=...)
# Second derivative d^2y/dx^2 = 6x
d2y = torch.autograd.grad(dy, x)[0]
print(d2y) # tensor(6.)Autograd in the Training Pipeline
Autograd is the foundation of every neural network training loop. The standard four-step pattern is: (1) zero gradients with optimizer.zero_grad(), (2) forward pass to compute predictions and loss, (3) backward pass with loss.backward() to compute gradients, and (4) optimizer step with optimizer.step() to update parameters. Understanding that autograd is doing the heavy calculus work lets you focus on model architecture and hyperparameters.
import torch
import torch.nn as nn
import torch.optim as optim
model = nn.Linear(1, 1)
optimizer = optim.SGD(model.parameters(), lr=0.1)
criterion = nn.MSELoss()
# Synthetic data: y = 3x + 1
X = torch.randn(20, 1)
y = 3 * X + 1 + 0.1 * torch.randn(20, 1)
for epoch in range(5):
optimizer.zero_grad() # (1) zero grads
y_pred = model(X) # (2) forward
loss = criterion(y_pred, y) # (2) loss
loss.backward() # (3) backward
optimizer.step() # (4) update
print(f'Epoch {epoch}: loss={loss.item():.4f}')Quick Check
Test your understanding of Machine Learning with Python concepts from this lesson.
Lesson Recap
In this lesson you learned: autograd builds a dynamic computation graph recording every operation on tensors with requires_grad=True, .backward() computes gradients via the chain rule and deposits them in .grad, and gradients accumulate so you must zero them before each training step. Next up we explore building a feedforward network with nn.Module.
常见问题解答
「Autograd:用于反向传播的自动微分」课时是免费的吗?
是的 — 「Autograd:用于反向传播的自动微分」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Machine Learning Academy 课程的其余内容,请升级到 CoddyKit PRO。 Machine Learning Academy 课程共包含 4 节课。
「Autograd:用于反向传播的自动微分」这节课中我会学到什么?
您将定义一个标量计算图,调用 .backward(),并检查叶张量上的 .grad,以理解梯度如何在网络中流动。 你通过在浏览器中直接运行的动手代码来练习 Machine Learning Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Machine Learning Academy 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Machine Learning Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 2 节课,共 4 节。
「Autograd:用于反向传播的自动微分」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Machine Learning Academy 课中编写并运行代码吗?
能。每节 Machine Learning Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- PyTorch 张量:创建、运算与 GPU 传输
- Autograd:用于反向传播的自动微分
- 使用 nn.Module 构建前馈网络
- 训练循环:损失、优化器与训练轮次