0Pricing
Deep Learning Academy · درس

التدرجات المتلاشية والمتفجرة

ما الذي يعطّل الشبكات العميقة والإصلاحات الأولى

التدرجات المتلاشية والمتفجرة درس مجاني في Deep Learning Academy على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Deep Learning Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Deep Learning Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

Gradients Are a Product

Through many layers, backprop multiplies many local derivatives together. The size of that long product decides whether learning works or breaks.

Multiplying Small Numbers

If each link's derivative is below one, the product shrinks fast. After many layers the gradient becomes tiny, almost zero by the time it reaches early layers.

That Is Vanishing Gradients

When gradients shrink to near zero, early layers barely update and stop learning. This is the vanishing gradient problem that long stalled deep nets.

Multiplying Big Numbers

If each derivative is above one, the product blows up instead. Gradients grow huge as they travel back, and the weights lurch wildly.

That Is Exploding Gradients

Runaway gradients are the exploding gradient problem. Loss often jumps to NaN as updates overshoot far past any useful value. 💥

Sigmoid Made It Worse

Sigmoid and tanh squash inputs, so their derivatives stay well below one. Stacking them multiplied small numbers and made vanishing gradients common.

ReLU to the Rescue

ReLU has a derivative of exactly one for positive inputs, so it does not shrink the gradient. Switching to ReLU was a key early fix.

relu_grad = 1.0 if x > 0 else 0.0

Careful Weight Initialization

Smart schemes like Xavier and He initialization scale starting weights so the gradient product stays near one across many layers.

Clip Exploding Gradients

For explosions, gradient clipping caps the gradient's size before the update, keeping a single huge step from wrecking the model.

torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)

Skip Connections Help

ResNet adds skip connections that let gradients flow straight back, sidestepping the long chain of multiplications that causes vanishing.

Why It All Matters

Keeping the gradient product near one is what lets very deep nets train at all. Every fix here exists to protect that balance.

Quick Check

Let's check the failure modes.

Recap

Long products of derivatives can vanish or explode. ReLU, careful init, clipping, and skip connections keep gradients healthy through deep nets. 💥

الأسئلة الشائعة

هل درس «التدرجات المتلاشية والمتفجرة» مجاني؟

نعم — نص درس «التدرجات المتلاشية والمتفجرة» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Deep Learning Academy، انتقل إلى CoddyKit PRO. تتضمن دورة Deep Learning Academy 4 دروس في المجموع.

ماذا ستتعلم في «التدرجات المتلاشية والمتفجرة»؟

ما الذي يعطّل الشبكات العميقة والإصلاحات الأولى تتمرن على Deep Learning Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ Deep Learning Academy؟

لا تُشترط خبرة سابقة. Deep Learning Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.

كم من الوقت يستغرق درس «التدرجات المتلاشية والمتفجرة»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس Deep Learning Academy هذا؟

نعم. كل درس في Deep Learning Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. قاعدة السلسلة، طبقة بعد طبقة
  2. التخزين المؤقت في المرور الأمامي وإعادة الاستخدام في العكسي
  3. إجراء Backprop لشبكة صغيرة يدويًا
  4. التدرجات المتلاشية والمتفجرة
← العودة إلى Deep Learning Academy