0Pricing
Deep Learning Academy · Lección

Weight decay frente a regularización L2

La sutil diferencia que importa

Weight decay frente a regularización L2 es una lección gratuita de Deep Learning Academy en CoddyKit. Esta es la lección 3 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Deep Learning Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Deep Learning Academy incluye 4 lecciones en total.

Partes de esta lección aún no han sido traducidas y se muestran en inglés.

Keep Weights Small

Big weights often mean an overfit model. Both weight decay and L2 regularization push weights toward zero so the network stays simpler.

L2 Adds to the Loss

L2 regularization adds a penalty term, the sum of squared weights, straight into the loss. Minimizing loss then also means shrinking the weights.

loss = data_loss + lam * (w ** 2).sum()

Decay Shrinks Directly

Weight decay skips the loss and instead multiplies each weight by a number slightly under one every step. It shrinks weights before the gradient update.

w = w - lr * grad - lr * wd * w

Same Thing for SGD

With plain SGD, the two are mathematically identical. The L2 penalty's gradient is exactly the decay term, so it makes no practical difference.

Adam Breaks the Tie

The catch appears with Adam. Its per-weight scaling divides the L2 gradient unevenly, so L2 and true weight decay stop being equal.

Why It Distorts

Adam shrinks weights with large gradients less and small ones more. Folded-in L2 inherits that bias, weakening the regularization where you need it.

Decoupling Is the Fix

AdamW applies decoupled weight decay, shrinking every weight by the same fraction independent of its gradient. That restores honest regularization.

Set It in PyTorch

The weight_decay argument controls the strength. In AdamW it is true decoupled decay, the behavior you usually want.

opt = torch.optim.AdamW(model.parameters(), lr=1e-3, weight_decay=0.01)

Pick a Strength

Values like 0.01 or 0.0001 are typical. Too much decay underfits; too little lets weights grow and overfit. Tune it like any hyperparameter.

Spare the Biases

It is common to skip decay on bias and norm parameters. Shrinking them rarely helps and can quietly hurt how the model trains.

The Takeaway

With SGD, reach for either name freely. With adaptive optimizers, prefer AdamW so your weight decay actually behaves as designed.

Quick Check

Test when the two truly diverge.

Recap

L2 adds a penalty to the loss; weight decay shrinks weights directly. They match under SGD but split under Adam, which is why AdamW decouples decay. 🪶

Preguntas frecuentes

¿La lección «Weight decay frente a regularización L2» es gratis?

Sí — el texto completo de «Weight decay frente a regularización L2» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Deep Learning Academy, actualiza a CoddyKit PRO. El curso de Deep Learning Academy incluye 4 lecciones en total.

¿Qué aprenderé en «Weight decay frente a regularización L2»?

La sutil diferencia que importa Practicas Deep Learning Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.

¿Necesito experiencia previa para empezar Deep Learning Academy?

No se requiere experiencia previa. Deep Learning Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 3 de 4.

¿Cuánto tiempo toma la lección «Weight decay frente a regularización L2»?

La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.

¿Puedo escribir y ejecutar código en esta lección de Deep Learning Academy?

Sí. Cada lección de Deep Learning Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.

Todas las lecciones de este curso

  1. SGD con momentum
  2. Adam y AdamW explicados
  3. Weight decay frente a regularización L2
  4. Programaciones de la tasa de aprendizaje y warmup
← Volver a Deep Learning Academy