Вентили LSTM и GRU
Запоминайте долгосрочные зависимости
«Вентили LSTM и GRU» — бесплатный урок Deep Learning Academy на CoddyKit. Это урок 3 из 4. Ты можешь прочитать весь урок бесплатно ниже — а потом практиковать его прямо в браузере с встроенным редактором кода и ИИ-репетитором 24/7. Это часть пути обучения Deep Learning Academy, и твой прогресс синхронизируется между веб-версией и приложением CoddyKit. Курс Deep Learning Academy содержит 4 уроков всего.
Части этого урока еще не переведены и отображаются на английском.
The Long-Memory Problem
Vanilla RNNs forget early clues over long sequences. Gated cells fix this by deciding what to keep, update, or throw away.
Meet the LSTM
The LSTM adds a separate cell state, a memory highway that runs straight across time with only small, controlled edits.
Gates Are Soft Switches
A gate is a sigmoid layer outputting values from 0 to 1. Zero blocks information, one lets it pass, and in between mixes the two.
gate = torch.sigmoid(W @ x + U @ h)The Forget Gate
The forget gate looks at the input and memory and chooses which parts of the old cell state to erase before adding anything new.
The Input Gate
The input gate decides how much of the fresh candidate information should be written into the cell state at this step.
The Output Gate
The output gate controls how much of the updated cell state becomes the visible hidden state passed to the next step. 🚪
Why Gates Help Gradients
Because the cell state changes gently, gradients flow back across many steps without vanishing, so the model learns long-range patterns.
Meet the GRU
The GRU is a lighter cousin: it merges gates and drops the separate cell state, giving similar power with fewer parameters.
GRU's Two Gates
A GRU uses just two gates, a reset gate and an update gate, to balance old memory against new input each step.
Drop-In in PyTorch
Both are one-liners in PyTorch. Swap nn.LSTM or nn.GRU for nn.RNN and keep almost the same training code.
lstm = nn.LSTM(input_size=10, hidden_size=20)Which to Choose?
GRUs are faster and often match LSTMs; LSTMs can edge ahead on the hardest long sequences. Try both and let your validation score decide.
Quick Check
What is the job of the forget gate in an LSTM?
Recap
LSTMs and GRUs use gates to control memory, letting gradients survive long sequences and capturing far-apart dependencies. ✅
Часто задаваемые вопросы
Урок «Вентили LSTM и GRU» бесплатный?
Да — полный текст урока «Вентили LSTM и GRU» бесплатно доступен здесь в веб-версии. Чтобы практиковать его интерактивно (встроенный редактор кода и ИИ-репетитор 24/7) и разблокировать остальной курс Deep Learning Academy, подпишись на CoddyKit PRO. Курс Deep Learning Academy содержит 4 уроков всего.
Чему я научусь в уроке «Вентили LSTM и GRU»?
Запоминайте долгосрочные зависимости Ты практикуешь Deep Learning Academy с помощью реального кода, который запускаешь прямо в браузере, и ИИ-репетитор 24/7 отвечает на твои вопросы во время урока.
Нужен ли мне опыт, чтобы начать Deep Learning Academy?
Предыдущий опыт не требуется. Deep Learning Academy на CoddyKit структурирован для всех уровней — от новичков до продвинутых, поэтому ты можешь начать отсюда или с самого начала и учиться в своем темпе. Это урок 3 из 4.
Сколько времени занимает урок «Вентили LSTM и GRU»?
Большинство уроков CoddyKit занимают около 5–10 минут. Каждый из них компактный и интерактивный, поэтому ты постоянно делаешь прогресс и продолжаешь с того же места в веб-версии и приложении.
Можно ли писать и запускать код в этом уроке Deep Learning Academy?
Да. Каждый урок Deep Learning Academy включает встроенный редактор кода, поэтому ты пишешь и запускаешь реальный код прямо в браузере и получаешь моментальную обратную связь от AI — локальная установка не требуется.
Все уроки этого курса
- Почему последовательностям нужна память
- Обычная ячейка RNN
- Вентили LSTM и GRU
- Упакуйте последовательности и обработайте дополнение