تجميع التسلسلات والتعامل مع Padding
درّب بكفاءة على أطوال متفاوتة
تجميع التسلسلات والتعامل مع Padding درس مجاني في Deep Learning Academy على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Deep Learning Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Deep Learning Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
Batches Want Rectangles
To process many sequences at once, a batch must be a neat rectangle. But real sentences have different lengths, so they don't line up.
Padding to a Common Length
The fix is padding: add filler tokens, usually zeros, to short sequences so every row in the batch is the same length.
pad_sequence Helper
PyTorch's pad_sequence stacks a list of variable-length tensors into one padded batch in a single call.
from torch.nn.utils.rnn import pad_sequence
batch = pad_sequence(seqs, batch_first=True)Padding Wastes Compute
Those filler tokens are fake. If the RNN processes them anyway, it burns compute and can let the padding pollute the hidden state.
Track the Real Lengths
Before padding, record each sequence's true length. You'll hand these lengths to PyTorch so it knows where the real data ends.
lengths = [len(s) for s in seqs]Packing the Batch
pack_padded_sequence turns the padded batch plus lengths into a compact form the RNN can run without touching the padding.
packed = pack_padded_sequence(batch, lengths, batch_first=True)Run the RNN on Packed Data
Feed the packed object straight into nn.LSTM or nn.GRU. The cell skips padded steps, so memory stays clean and clear.
out_packed, h_n = lstm(packed)Unpack the Output
To read per-step outputs again, call pad_packed_sequence to restore the rectangular padded shape after the RNN finishes.
out, lengths = pad_packed_sequence(out_packed, batch_first=True)Sort by Length
Classic packing wants sequences sorted longest first. Set enforce_sorted to False and PyTorch handles unsorted batches for you.
Mask the Loss
Even with packing, ignore padded positions when scoring. A mask or ignore_index keeps fake tokens out of the loss.
loss_fn = nn.CrossEntropyLoss(ignore_index=PAD_ID)Wire It Into the DataLoader
Do the padding inside a custom collate_fn so every batch arrives already padded, with lengths ready for packing.
Quick Check
Why pack a padded batch before feeding it to an RNN?
Recap
Pad to align a batch, track true lengths, then pack so the RNN ignores filler. Mask padded tokens in the loss too. ✅
الأسئلة الشائعة
هل درس «تجميع التسلسلات والتعامل مع Padding» مجاني؟
نعم — نص درس «تجميع التسلسلات والتعامل مع Padding» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Deep Learning Academy، انتقل إلى CoddyKit PRO. تتضمن دورة Deep Learning Academy 4 دروس في المجموع.
ماذا ستتعلم في «تجميع التسلسلات والتعامل مع Padding»؟
درّب بكفاءة على أطوال متفاوتة تتمرن على Deep Learning Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ Deep Learning Academy؟
لا تُشترط خبرة سابقة. Deep Learning Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.
كم من الوقت يستغرق درس «تجميع التسلسلات والتعامل مع Padding»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس Deep Learning Academy هذا؟
نعم. كل درس في Deep Learning Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- لماذا تحتاج التسلسلات إلى ذاكرة
- خلية RNN التقليدية
- بوابات LSTM وGRU
- تجميع التسلسلات والتعامل مع Padding