0Pricing
NLP Academy · درس

الانتباه متعدد الرؤوس والمواقع

وجهات نظر متعددة مع الوعي بالترتيب

الانتباه متعدد الرؤوس والمواقع درس مجاني في NLP Academy على CoddyKit. هذا هو الدرس 3 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في NLP Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة NLP Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

One Head Is Limiting

A single attention head can track only one kind of relationship at a time. Real language needs many patterns noticed at once. 🧠

Many Heads, Many Views

Multi-head attention runs several attention computations in parallel, each with its own learned projections and its own focus.

What Each Head Learns

One head might track subject-verb links while another follows pronouns. Together they capture a far richer picture of the sentence.

Split the Dimensions

The model splits its vector size across heads, so each head works in a smaller subspace. The total compute stays roughly the same.

d_head = d_model // num_heads

Combine the Heads

After each head produces an output, you concatenate them and pass the result through one more linear layer to mix the views.

out = concat(head_1, head_2, ...) @ W_o

Attention Ignores Order

Self-attention treats input as a set, so by itself it cannot tell "dog bites man" from "man bites dog." It is order-blind.

Adding Position Information

To fix this, we inject a positional encoding into each word so the model knows where every token sits in the sequence.

Sinusoidal Encodings

The original Transformer uses fixed sine and cosine waves of different frequencies to give each position a unique, smooth signature.

pe[pos, 2i] = sin(pos / 10000 ** (2*i/d))

Added, Not Appended

Position vectors are added to the word embeddings, not stuck on the end. So each token carries both meaning and place together.

x = token_embeddings + positional_encoding

Learned Positions Too

Many modern models replace fixed waves with learned position embeddings, trained alongside everything else for flexibility.

Why Both Matter

Multi-head attention sees many relationships; positional encodings restore order. Together they let the Transformer truly understand sequences. ✨

Quick Check

Let us test positions and heads.

Recap

You saw how multi-head attention captures many relationships in parallel, while positional encodings give the model a sense of order. 🎯

الأسئلة الشائعة

هل درس «الانتباه متعدد الرؤوس والمواقع» مجاني؟

نعم — نص درس «الانتباه متعدد الرؤوس والمواقع» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة NLP Academy، انتقل إلى CoddyKit PRO. تتضمن دورة NLP Academy 4 دروس في المجموع.

ماذا ستتعلم في «الانتباه متعدد الرؤوس والمواقع»؟

وجهات نظر متعددة مع الوعي بالترتيب تتمرن على NLP Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ NLP Academy؟

لا تُشترط خبرة سابقة. NLP Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 3 من أصل 4.

كم من الوقت يستغرق درس «الانتباه متعدد الرؤوس والمواقع»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس NLP Academy هذا؟

نعم. كل درس في NLP Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. فكرة الانتباه
  2. الانتباه الذاتي خطوة بخطوة
  3. الانتباه متعدد الرؤوس والمواقع
  4. داخل كتلة Transformer
← العودة إلى NLP Academy