0Pricing
Data Science Academy · درس

التعويض بالمتوسط أو الوسيط أو المنوال

ملء مناسب حسب نوع العمود

التعويض بالمتوسط أو الوسيط أو المنوال درس مجاني في Data Science Academy على CoddyKit. هذا هو الدرس 3 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Data Science Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Data Science Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

Filling With a Statistic

Instead of inventing random numbers, you fill gaps with a summary value from the column itself. This careful approach is called imputation. 🧮

The Mean Fill

For numbers, the simplest choice is the mean, the column average. It keeps the overall total roughly intact when gaps are few.

df['age'].fillna(df['age'].mean())

When the Mean Misleads

The mean is fragile: a few huge values drag it far from the center. On skewed data, mean imputation can pull every gap toward an unrealistic number.

The Median Fill

The median is the middle value, and it shrugs off extreme outliers. For skewed numeric columns it is usually the safer fill.

df['income'].fillna(df['income'].median())

Mean or Median?

A quick rule: reach for the median when a column has outliers or a long tail, and the mean only when values are fairly symmetric.

The Mode for Categories

Text and category columns have no average. For them you fill with the mode, the most frequent value, since that is the most likely fit.

top = df['city'].mode()[0]
df['city'].fillna(top)

Why mode Returns a Series

A column can tie for most common, so mode() returns a Series of all winners. Pick the first with index 0 when you need one value.

df['city'].mode()[0]

Fill Per Column Type

The best practice is to impute each column by its type: median for skewed numbers, mean for symmetric ones, and mode for categories.

Forward and Back Fill

For ordered data like time series, carry the last known value forward with ffill, or pull the next one backward with bfill.

df['temp'].ffill()

The Hidden Cost

Every imputation shrinks the column's natural spread, because filled values cluster at one point. Note this, since it nudges your variance downward.

Flag What You Filled

A pro habit: add a boolean column marking which rows were imputed. That flag lets later analysis know which values were real and which were guessed.

df['age_filled'] = df['age'].isna()

Quick Check

An income column is heavily skewed by a few millionaires. Which imputation fits best?

Recap: Smart Fills

You now impute with mean, median, or mode by column type, use ffill for ordered data, and flag filled rows so nothing gets hidden.

الأسئلة الشائعة

هل درس «التعويض بالمتوسط أو الوسيط أو المنوال» مجاني؟

نعم — نص درس «التعويض بالمتوسط أو الوسيط أو المنوال» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Data Science Academy، انتقل إلى CoddyKit PRO. تتضمن دورة Data Science Academy 4 دروس في المجموع.

ماذا ستتعلم في «التعويض بالمتوسط أو الوسيط أو المنوال»؟

ملء مناسب حسب نوع العمود تتمرن على Data Science Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ Data Science Academy؟

لا تُشترط خبرة سابقة. Data Science Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 3 من أصل 4.

كم من الوقت يستغرق درس «التعويض بالمتوسط أو الوسيط أو المنوال»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس Data Science Academy هذا؟

نعم. كل درس في Data Science Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. العثور على NaNs المختبئة في جدولك
  2. الحذف أم الملء: الاختيار بحكمة
  3. التعويض بالمتوسط أو الوسيط أو المنوال
  4. إصلاح dtypes والصفوف المكررة
← العودة إلى Data Science Academy