Leer y poner a cero .grad
Por qué los gradientes se acumulan y deben reiniciarse
Leer y poner a cero .grad es una lección gratuita de Deep Learning Academy en CoddyKit. Esta es la lección 3 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Deep Learning Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Deep Learning Academy incluye 4 lecciones en total.
Partes de esta lección aún no han sido traducidas y se muestran en inglés.
The .grad Attribute
After backward, every trained tensor stores its gradient in .grad. Reading it tells you how the loss responds to changes in that tensor.
print(weight.grad)Gradients Start as None
Before any backward call, .grad is None, not zero. PyTorch only allocates it once the first set of gradients actually arrives.
Gradients Accumulate
Here is the surprise: backward adds to .grad rather than replacing it. Call backward twice without clearing and the numbers pile up.
Why Accumulation Exists
This adding behavior is on purpose. It lets you sum gradients from several mini-batches before one update, handy for simulating a larger batch.
The Hidden Bug
Forget to clear and old gradients poison your next step, so the model trains on the wrong numbers. This silent bug trips up almost everyone once.
Zero Them Each Step
The fix is to reset gradients before each backward. With an optimizer you simply call zero_grad() at the top of every training step.
optimizer.zero_grad()The Right Order
The loop rhythm is fixed: zero_grad, forward, loss, backward, step. Zeroing first guarantees each step uses only this batch's gradients.
optimizer.zero_grad()
loss.backward()
optimizer.step()Clearing Without an Optimizer
No optimizer yet? You can null the gradients yourself by setting each tensor's .grad back to None before the next backward call.
w.grad = Noneset_to_none Is Cheaper
Newer code prefers zero_grad(set_to_none=True). Setting grads to None instead of filling zeros saves a little memory and time.
Reading Grads to Debug
Peeking at .grad is great for debugging. All zeros may mean a dead neuron, and huge values warn of exploding gradients before training blows up.
A Habit Worth Forming
Make zeroing automatic in your head. Every reliable training loop clears gradients first, so the model only ever learns from the current batch.
Quick Check
Recap
Gradients live in .grad and accumulate across backward calls, so you must clear them each step with zero_grad. Forgetting that quietly breaks training. 🧹
Preguntas frecuentes
¿La lección «Leer y poner a cero .grad» es gratis?
Sí — el texto completo de «Leer y poner a cero .grad» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Deep Learning Academy, actualiza a CoddyKit PRO. El curso de Deep Learning Academy incluye 4 lecciones en total.
¿Qué aprenderé en «Leer y poner a cero .grad»?
Por qué los gradientes se acumulan y deben reiniciarse Practicas Deep Learning Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.
¿Necesito experiencia previa para empezar Deep Learning Academy?
No se requiere experiencia previa. Deep Learning Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 3 de 4.
¿Cuánto tiempo toma la lección «Leer y poner a cero .grad»?
La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.
¿Puedo escribir y ejecutar código en esta lección de Deep Learning Academy?
Sí. Cada lección de Deep Learning Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.
Todas las lecciones de este curso
- requires_grad y el grafo de cómputo
- Llame a backward() para obtener gradientes
- Leer y poner a cero .grad
- torch.no_grad() para inferencia