Adam y AdamW explicados
Tasas adaptativas y weight decay desacoplado
Adam y AdamW explicados es una lección gratuita de Deep Learning Academy en CoddyKit. Esta es la lección 2 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Deep Learning Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Deep Learning Academy incluye 4 lecciones en total.
Partes de esta lección aún no han sido traducidas y se muestran en inglés.
One Rate Per Weight
SGD uses a single learning rate for every parameter. Adam adapts the step size for each weight on its own, based on that weight's gradient history.
Two Moving Averages
Adam tracks two running averages: the mean of gradients and the mean of their squares. Together they form the first and second moments.
First Moment Is Momentum
The first moment is basically momentum, the smoothed average of recent gradients. It decides the overall direction each weight should move.
Second Moment Scales Steps
The second moment estimates each gradient's size. Adam divides by its square root, so noisy weights take smaller steps and quiet ones take larger.
The Betas
Two decay rates, the betas, control those averages, typically 0.9 and 0.999. They balance how much recent versus older gradients matter.
Bias Correction
The averages start at zero, so early steps look too small. Adam applies a bias correction to fix this so updates are sane from step one.
Use It in PyTorch
One line gives you Adam. The default learning rate of 0.001 works well across a huge range of models, which is why it is so popular.
opt = torch.optim.Adam(model.parameters(), lr=1e-3)Adam's Weight Decay Flaw
Classic Adam mixes weight decay into the gradient, where the adaptive scaling distorts it. The regularization ends up weaker than you intended.
AdamW Fixes It
AdamW decouples weight decay from the gradient step and applies it directly to the weights. The decay now works as a clean, predictable shrink.
opt = torch.optim.AdamW(model.parameters(), lr=1e-3, weight_decay=0.01)The Modern Default
For transformers and most large models, AdamW is the standard choice. Reach for it first whenever you want fast, reliable convergence.
Adaptive, With Caveats
Adam often trains faster than SGD, yet plain SGD with momentum can generalize better on vision tasks. Try both when accuracy really counts.
Quick Check
Pin down the Adam to AdamW difference.
Recap
Adam adapts a learning rate per weight using gradient mean and variance, while AdamW fixes its weight decay. AdamW is today's go-to optimizer. ⚙️
Preguntas frecuentes
¿La lección «Adam y AdamW explicados» es gratis?
Sí — el texto completo de «Adam y AdamW explicados» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Deep Learning Academy, actualiza a CoddyKit PRO. El curso de Deep Learning Academy incluye 4 lecciones en total.
¿Qué aprenderé en «Adam y AdamW explicados»?
Tasas adaptativas y weight decay desacoplado Practicas Deep Learning Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.
¿Necesito experiencia previa para empezar Deep Learning Academy?
No se requiere experiencia previa. Deep Learning Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 2 de 4.
¿Cuánto tiempo toma la lección «Adam y AdamW explicados»?
La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.
¿Puedo escribir y ejecutar código en esta lección de Deep Learning Academy?
Sí. Cada lección de Deep Learning Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.
Todas las lecciones de este curso
- SGD con momentum
- Adam y AdamW explicados
- Weight decay frente a regularización L2
- Programaciones de la tasa de aprendizaje y warmup