SGD com Momentum
Suavize as atualizações para superar o ruído
SGD com Momentum é uma aula grátis de Deep Learning Academy no CoddyKit. Esta é a aula 1 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de Deep Learning Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de Deep Learning Academy inclui 4 aulas no total.
Partes desta aula ainda não foram traduzidas e aparecem em inglês.
Plain SGD Forgets
Vanilla SGD steps using only the current gradient. Every update starts from scratch, so a noisy slope makes it wobble and crawl toward the minimum.
Borrow Some Inertia
Momentum gives SGD memory. It keeps a running average of past gradients and rolls in that direction, like a ball gathering speed downhill.
The Velocity Vector
Momentum tracks a velocity that blends the old velocity with the new gradient. The weights then move by that smoothed velocity each step.
v = beta * v + grad
w = w - lr * vThe Beta Knob
The momentum coefficient beta sets how much past steps count. A common value is 0.9, meaning most of the velocity carries over.
Smooths Out Noise
Because momentum averages many gradients, random noise in any single batch mostly cancels. The path to the minimum becomes smoother and steadier.
Rolls Past Small Bumps
Built-up speed lets the optimizer coast through tiny dips and flat spots that would stall plain SGD. The ball does not stop at every pebble.
Turn It On in PyTorch
You do not code this by hand. Just pass momentum to the built-in SGD optimizer and PyTorch tracks the velocity for you.
opt = torch.optim.SGD(model.parameters(), lr=0.01, momentum=0.9)Nesterov Looks Ahead
A sharper variant, Nesterov momentum, peeks at where the velocity is heading before measuring the gradient. It often converges a touch faster.
opt = torch.optim.SGD(model.parameters(), lr=0.01, momentum=0.9, nesterov=True)Mind the Overshoot
Too much speed can carry you past the valley floor. If loss bounces or diverges, lower the learning rate or trim momentum a little.
Why It Still Matters
Even with fancy optimizers around, SGD with momentum remains a strong baseline and often generalizes beautifully on vision models.
A Faster Descent
The payoff is real: momentum usually reaches a good minimum in fewer epochs than plain SGD, with less zig-zagging along the way.
Quick Check
Make sure momentum's role is clear.
Recap
Momentum gives SGD a velocity that blends past gradients, smoothing noise and coasting past small bumps. Set momentum near 0.9 for a faster, steadier descent. 🏂
Perguntas Frequentes
A aula “SGD com Momentum” é grátis?
Sim — o texto completo de “SGD com Momentum” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de Deep Learning Academy, atualize para CoddyKit PRO. O curso de Deep Learning Academy inclui 4 aulas no total.
O que vou aprender em “SGD com Momentum”?
Suavize as atualizações para superar o ruído Você pratica Deep Learning Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.
Preciso ter experiência prévia para começar Deep Learning Academy?
Nenhuma experiência prévia é necessária. Deep Learning Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 1 de 4.
Quanto tempo leva a aula “SGD com Momentum”?
A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.
Posso escrever e executar código nesta aula de Deep Learning Academy?
Sim. Cada aula de Deep Learning Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.
Todas as aulas deste curso
- SGD com Momentum
- Adam e AdamW Explicados
- Decaimento de Pesos vs Regularização L2
- Agendamentos da Taxa de Aprendizado e Aquecimento