0Pricing
Deep Learning Academy · Aula

Decaimento de Pesos vs Regularização L2

A diferença sutil que importa

Decaimento de Pesos vs Regularização L2 é uma aula grátis de Deep Learning Academy no CoddyKit. Esta é a aula 3 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de Deep Learning Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de Deep Learning Academy inclui 4 aulas no total.

Partes desta aula ainda não foram traduzidas e aparecem em inglês.

Keep Weights Small

Big weights often mean an overfit model. Both weight decay and L2 regularization push weights toward zero so the network stays simpler.

L2 Adds to the Loss

L2 regularization adds a penalty term, the sum of squared weights, straight into the loss. Minimizing loss then also means shrinking the weights.

loss = data_loss + lam * (w ** 2).sum()

Decay Shrinks Directly

Weight decay skips the loss and instead multiplies each weight by a number slightly under one every step. It shrinks weights before the gradient update.

w = w - lr * grad - lr * wd * w

Same Thing for SGD

With plain SGD, the two are mathematically identical. The L2 penalty's gradient is exactly the decay term, so it makes no practical difference.

Adam Breaks the Tie

The catch appears with Adam. Its per-weight scaling divides the L2 gradient unevenly, so L2 and true weight decay stop being equal.

Why It Distorts

Adam shrinks weights with large gradients less and small ones more. Folded-in L2 inherits that bias, weakening the regularization where you need it.

Decoupling Is the Fix

AdamW applies decoupled weight decay, shrinking every weight by the same fraction independent of its gradient. That restores honest regularization.

Set It in PyTorch

The weight_decay argument controls the strength. In AdamW it is true decoupled decay, the behavior you usually want.

opt = torch.optim.AdamW(model.parameters(), lr=1e-3, weight_decay=0.01)

Pick a Strength

Values like 0.01 or 0.0001 are typical. Too much decay underfits; too little lets weights grow and overfit. Tune it like any hyperparameter.

Spare the Biases

It is common to skip decay on bias and norm parameters. Shrinking them rarely helps and can quietly hurt how the model trains.

The Takeaway

With SGD, reach for either name freely. With adaptive optimizers, prefer AdamW so your weight decay actually behaves as designed.

Quick Check

Test when the two truly diverge.

Recap

L2 adds a penalty to the loss; weight decay shrinks weights directly. They match under SGD but split under Adam, which is why AdamW decouples decay. 🪶

Perguntas Frequentes

A aula “Decaimento de Pesos vs Regularização L2” é grátis?

Sim — o texto completo de “Decaimento de Pesos vs Regularização L2” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de Deep Learning Academy, atualize para CoddyKit PRO. O curso de Deep Learning Academy inclui 4 aulas no total.

O que vou aprender em “Decaimento de Pesos vs Regularização L2”?

A diferença sutil que importa Você pratica Deep Learning Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.

Preciso ter experiência prévia para começar Deep Learning Academy?

Nenhuma experiência prévia é necessária. Deep Learning Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 3 de 4.

Quanto tempo leva a aula “Decaimento de Pesos vs Regularização L2”?

A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.

Posso escrever e executar código nesta aula de Deep Learning Academy?

Sim. Cada aula de Deep Learning Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.

Todas as aulas deste curso

  1. SGD com Momentum
  2. Adam e AdamW Explicados
  3. Decaimento de Pesos vs Regularização L2
  4. Agendamentos da Taxa de Aprendizado e Aquecimento
← Voltar para Deep Learning Academy