0Pricing
Deep Learning Academy · Aula

SGD com Momentum

Suavize as atualizações para superar o ruído

SGD com Momentum é uma aula grátis de Deep Learning Academy no CoddyKit. Esta é a aula 1 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de Deep Learning Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de Deep Learning Academy inclui 4 aulas no total.

Partes desta aula ainda não foram traduzidas e aparecem em inglês.

Plain SGD Forgets

Vanilla SGD steps using only the current gradient. Every update starts from scratch, so a noisy slope makes it wobble and crawl toward the minimum.

Borrow Some Inertia

Momentum gives SGD memory. It keeps a running average of past gradients and rolls in that direction, like a ball gathering speed downhill.

The Velocity Vector

Momentum tracks a velocity that blends the old velocity with the new gradient. The weights then move by that smoothed velocity each step.

v = beta * v + grad
w = w - lr * v

The Beta Knob

The momentum coefficient beta sets how much past steps count. A common value is 0.9, meaning most of the velocity carries over.

Smooths Out Noise

Because momentum averages many gradients, random noise in any single batch mostly cancels. The path to the minimum becomes smoother and steadier.

Rolls Past Small Bumps

Built-up speed lets the optimizer coast through tiny dips and flat spots that would stall plain SGD. The ball does not stop at every pebble.

Turn It On in PyTorch

You do not code this by hand. Just pass momentum to the built-in SGD optimizer and PyTorch tracks the velocity for you.

opt = torch.optim.SGD(model.parameters(), lr=0.01, momentum=0.9)

Nesterov Looks Ahead

A sharper variant, Nesterov momentum, peeks at where the velocity is heading before measuring the gradient. It often converges a touch faster.

opt = torch.optim.SGD(model.parameters(), lr=0.01, momentum=0.9, nesterov=True)

Mind the Overshoot

Too much speed can carry you past the valley floor. If loss bounces or diverges, lower the learning rate or trim momentum a little.

Why It Still Matters

Even with fancy optimizers around, SGD with momentum remains a strong baseline and often generalizes beautifully on vision models.

A Faster Descent

The payoff is real: momentum usually reaches a good minimum in fewer epochs than plain SGD, with less zig-zagging along the way.

Quick Check

Make sure momentum's role is clear.

Recap

Momentum gives SGD a velocity that blends past gradients, smoothing noise and coasting past small bumps. Set momentum near 0.9 for a faster, steadier descent. 🏂

Perguntas Frequentes

A aula “SGD com Momentum” é grátis?

Sim — o texto completo de “SGD com Momentum” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de Deep Learning Academy, atualize para CoddyKit PRO. O curso de Deep Learning Academy inclui 4 aulas no total.

O que vou aprender em “SGD com Momentum”?

Suavize as atualizações para superar o ruído Você pratica Deep Learning Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.

Preciso ter experiência prévia para começar Deep Learning Academy?

Nenhuma experiência prévia é necessária. Deep Learning Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 1 de 4.

Quanto tempo leva a aula “SGD com Momentum”?

A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.

Posso escrever e executar código nesta aula de Deep Learning Academy?

Sim. Cada aula de Deep Learning Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.

Todas as aulas deste curso

  1. SGD com Momentum
  2. Adam e AdamW Explicados
  3. Decaimento de Pesos vs Regularização L2
  4. Agendamentos da Taxa de Aprendizado e Aquecimento
← Voltar para Deep Learning Academy