SGD с моментом
Сглаживайте обновления, чтобы преодолевать шум
«SGD с моментом» — бесплатный урок Deep Learning Academy на CoddyKit. Это урок 1 из 4. Ты можешь прочитать весь урок бесплатно ниже — а потом практиковать его прямо в браузере с встроенным редактором кода и ИИ-репетитором 24/7. Это часть пути обучения Deep Learning Academy, и твой прогресс синхронизируется между веб-версией и приложением CoddyKit. Курс Deep Learning Academy содержит 4 уроков всего.
Части этого урока еще не переведены и отображаются на английском.
Plain SGD Forgets
Vanilla SGD steps using only the current gradient. Every update starts from scratch, so a noisy slope makes it wobble and crawl toward the minimum.
Borrow Some Inertia
Momentum gives SGD memory. It keeps a running average of past gradients and rolls in that direction, like a ball gathering speed downhill.
The Velocity Vector
Momentum tracks a velocity that blends the old velocity with the new gradient. The weights then move by that smoothed velocity each step.
v = beta * v + grad
w = w - lr * vThe Beta Knob
The momentum coefficient beta sets how much past steps count. A common value is 0.9, meaning most of the velocity carries over.
Smooths Out Noise
Because momentum averages many gradients, random noise in any single batch mostly cancels. The path to the minimum becomes smoother and steadier.
Rolls Past Small Bumps
Built-up speed lets the optimizer coast through tiny dips and flat spots that would stall plain SGD. The ball does not stop at every pebble.
Turn It On in PyTorch
You do not code this by hand. Just pass momentum to the built-in SGD optimizer and PyTorch tracks the velocity for you.
opt = torch.optim.SGD(model.parameters(), lr=0.01, momentum=0.9)Nesterov Looks Ahead
A sharper variant, Nesterov momentum, peeks at where the velocity is heading before measuring the gradient. It often converges a touch faster.
opt = torch.optim.SGD(model.parameters(), lr=0.01, momentum=0.9, nesterov=True)Mind the Overshoot
Too much speed can carry you past the valley floor. If loss bounces or diverges, lower the learning rate or trim momentum a little.
Why It Still Matters
Even with fancy optimizers around, SGD with momentum remains a strong baseline and often generalizes beautifully on vision models.
A Faster Descent
The payoff is real: momentum usually reaches a good minimum in fewer epochs than plain SGD, with less zig-zagging along the way.
Quick Check
Make sure momentum's role is clear.
Recap
Momentum gives SGD a velocity that blends past gradients, smoothing noise and coasting past small bumps. Set momentum near 0.9 for a faster, steadier descent. 🏂
Изучай Python с ИИ-репетитором — бесплатно
Пиши и запускай код прямо в браузере, получай мгновенную помощь от ИИ-репетитора 24/7 и продолжи учиться на сайте или в приложении.
- Курсы
- 30
- Уроки
- 120
Часто задаваемые вопросы
Урок «SGD с моментом» бесплатный?
Да — полный текст урока «SGD с моментом» бесплатно доступен здесь в веб-версии. Чтобы практиковать его интерактивно (встроенный редактор кода и ИИ-репетитор 24/7) и разблокировать остальной курс Deep Learning Academy, подпишись на CoddyKit PRO. Курс Deep Learning Academy содержит 4 уроков всего.
Чему я научусь в уроке «SGD с моментом»?
Сглаживайте обновления, чтобы преодолевать шум Ты практикуешь Deep Learning Academy с помощью реального кода, который запускаешь прямо в браузере, и ИИ-репетитор 24/7 отвечает на твои вопросы во время урока.
Нужен ли мне опыт, чтобы начать Deep Learning Academy?
Предыдущий опыт не требуется. Deep Learning Academy на CoddyKit структурирован для всех уровней — от новичков до продвинутых, поэтому ты можешь начать отсюда или с самого начала и учиться в своем темпе. Это урок 1 из 4.
Сколько времени занимает урок «SGD с моментом»?
Большинство уроков CoddyKit занимают около 5–10 минут. Каждый из них компактный и интерактивный, поэтому ты постоянно делаешь прогресс и продолжаешь с того же места в веб-версии и приложении.
Можно ли писать и запускать код в этом уроке Deep Learning Academy?
Да. Каждый урок Deep Learning Academy включает встроенный редактор кода, поэтому ты пишешь и запускаешь реальный код прямо в браузере и получаешь моментальную обратную связь от AI — локальная установка не требуется.
Все уроки этого курса
- SGD с моментом
- Adam и AdamW: объяснение
- Уменьшение весов и L2-регуляризация
- Расписания скорости обучения и разогрев