0Pricing
Deep Learning Academy · درس

الترميز الموضعي للترتيب

أدخل موضع التسلسل في Tokens

الترميز الموضعي للترتيب درس مجاني في Deep Learning Academy على CoddyKit. هذا هو الدرس 3 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Deep Learning Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Deep Learning Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

Attention Ignores Order

Self-attention treats a sentence like a bag of tokens. Shuffle the words and the math barely changes, so the model has no built-in sense of order.

Order Carries Meaning

But order matters: "dog bites man" is not "man bites dog". We must hand the model some signal about each token's position.

Add, Do Not Append

The trick is to add a position vector directly to each token's embedding. Same shape, so attention sees content and place fused together.

x = token_emb + pos_emb

Sinusoidal Encoding

The original transformer used fixed sine and cosine waves of different frequencies, one pattern per dimension, to mark each position.

Why Sine Waves

Mixing frequencies gives every position a unique fingerprint, and the smooth waves let the model generalize to lengths it never saw in training.

The Formula

Even dimensions use sine, odd dimensions use cosine, with the wavelength growing across dimensions. That is the whole encoding recipe.

pe[:, 0::2] = torch.sin(pos / div)
pe[:, 1::2] = torch.cos(pos / div)

Relative Distance

A neat bonus: with sinusoids, the offset between two positions is easy to express, helping attention reason about relative distance.

Learned Positions

Modern models often skip sinusoids and use a learned embedding table, one trainable vector per position, just like word embeddings.

self.pos_emb = nn.Embedding(max_len, d_model)

Fixed vs Learned

Fixed sinusoids extrapolate to new lengths for free; learned tables fit your data but are capped at the maximum length you trained on.

Rotary Encoding

Newer transformers favor rotary position encoding, which rotates query and key vectors by an angle tied to position, baking order into attention itself.

Where It Goes

Whatever scheme you pick, positional info is injected once at the input, before the first attention layer ever runs.

Quick Check

Let's check why positional encoding exists.

Recap

You learned that attention ignores order, so we add positional encodings, whether sinusoidal, learned, or rotary, to tell tokens where they sit. Well done!

الأسئلة الشائعة

هل درس «الترميز الموضعي للترتيب» مجاني؟

نعم — نص درس «الترميز الموضعي للترتيب» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Deep Learning Academy، انتقل إلى CoddyKit PRO. تتضمن دورة Deep Learning Academy 4 دروس في المجموع.

ماذا ستتعلم في «الترميز الموضعي للترتيب»؟

أدخل موضع التسلسل في Tokens تتمرن على Deep Learning Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ Deep Learning Academy؟

لا تُشترط خبرة سابقة. Deep Learning Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 3 من أصل 4.

كم من الوقت يستغرق درس «الترميز الموضعي للترتيب»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس Deep Learning Academy هذا؟

نعم. كل درس في Deep Learning Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. الانتباه الذاتي: Query وKey وValue
  2. الضرب النقطي المُقاس ومتعدد الرؤوس
  3. الترميز الموضعي للترتيب
  4. تكديس كتلة Transformer Encoder
← العودة إلى Deep Learning Academy