Empaquete secuencias y gestione el padding
Entrene eficazmente con longitudes variables
Empaquete secuencias y gestione el padding es una lección gratuita de Deep Learning Academy en CoddyKit. Esta es la lección 4 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Deep Learning Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Deep Learning Academy incluye 4 lecciones en total.
Partes de esta lección aún no han sido traducidas y se muestran en inglés.
Batches Want Rectangles
To process many sequences at once, a batch must be a neat rectangle. But real sentences have different lengths, so they don't line up.
Padding to a Common Length
The fix is padding: add filler tokens, usually zeros, to short sequences so every row in the batch is the same length.
pad_sequence Helper
PyTorch's pad_sequence stacks a list of variable-length tensors into one padded batch in a single call.
from torch.nn.utils.rnn import pad_sequence
batch = pad_sequence(seqs, batch_first=True)Padding Wastes Compute
Those filler tokens are fake. If the RNN processes them anyway, it burns compute and can let the padding pollute the hidden state.
Track the Real Lengths
Before padding, record each sequence's true length. You'll hand these lengths to PyTorch so it knows where the real data ends.
lengths = [len(s) for s in seqs]Packing the Batch
pack_padded_sequence turns the padded batch plus lengths into a compact form the RNN can run without touching the padding.
packed = pack_padded_sequence(batch, lengths, batch_first=True)Run the RNN on Packed Data
Feed the packed object straight into nn.LSTM or nn.GRU. The cell skips padded steps, so memory stays clean and clear.
out_packed, h_n = lstm(packed)Unpack the Output
To read per-step outputs again, call pad_packed_sequence to restore the rectangular padded shape after the RNN finishes.
out, lengths = pad_packed_sequence(out_packed, batch_first=True)Sort by Length
Classic packing wants sequences sorted longest first. Set enforce_sorted to False and PyTorch handles unsorted batches for you.
Mask the Loss
Even with packing, ignore padded positions when scoring. A mask or ignore_index keeps fake tokens out of the loss.
loss_fn = nn.CrossEntropyLoss(ignore_index=PAD_ID)Wire It Into the DataLoader
Do the padding inside a custom collate_fn so every batch arrives already padded, with lengths ready for packing.
Quick Check
Why pack a padded batch before feeding it to an RNN?
Recap
Pad to align a batch, track true lengths, then pack so the RNN ignores filler. Mask padded tokens in the loss too. ✅
Preguntas frecuentes
¿La lección «Empaquete secuencias y gestione el padding» es gratis?
Sí — el texto completo de «Empaquete secuencias y gestione el padding» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Deep Learning Academy, actualiza a CoddyKit PRO. El curso de Deep Learning Academy incluye 4 lecciones en total.
¿Qué aprenderé en «Empaquete secuencias y gestione el padding»?
Entrene eficazmente con longitudes variables Practicas Deep Learning Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.
¿Necesito experiencia previa para empezar Deep Learning Academy?
No se requiere experiencia previa. Deep Learning Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 4 de 4.
¿Cuánto tiempo toma la lección «Empaquete secuencias y gestione el padding»?
La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.
¿Puedo escribir y ejecutar código en esta lección de Deep Learning Academy?
Sí. Cada lección de Deep Learning Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.
Todas las lecciones de este curso
- Por qué las secuencias necesitan memoria
- La celda RNN básica
- Puertas LSTM y GRU
- Empaquete secuencias y gestione el padding