تحجيم الأرقام وتطبيعها
وضع الميزات على مقياس عادل
تحجيم الأرقام وتطبيعها درس مجاني في Data Science Academy على CoddyKit. هذا هو الدرس 3 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Data Science Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Data Science Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
Why Scale Numbers
Age ranges 0 to 90 while income ranges to millions. Without scaling, the bigger numbers dominate and quietly drown out the smaller ones. ⚖️
Distance-Based Models Care
Models that measure distance, like k-NN and k-means, are very sensitive to scale. Unscaled features hand all the power to the largest column.
Standardization
Standardization rescales a column to mean 0 and standard deviation 1. Values become how many standard deviations they sit from the average.
StandardScaler
The StandardScaler applies standardization for you. Fit it on your data, then transform any column into z-scores.
from sklearn.preprocessing import StandardScaler
z = StandardScaler().fit_transform(df[['age']])Normalization
Normalization squeezes values into a fixed range, usually 0 to 1. The smallest value maps to 0 and the largest to 1.
MinMaxScaler
Reach for MinMaxScaler to rescale into 0 to 1. It keeps the shape of your data while bounding the range.
from sklearn.preprocessing import MinMaxScaler
m = MinMaxScaler().fit_transform(df[['age']])Standardize or Normalize
Use standardization when data is roughly bell-shaped or has outliers. Pick normalization when you need a strict 0 to 1 bound.
Outliers Hurt MinMax
A single huge value can squash everything else near zero under MinMax. When outliers dominate, RobustScaler resists them better.
from sklearn.preprocessing import RobustScaler
r = RobustScaler().fit_transform(df[['income']])Fit on Train Only
Fit the scaler on training data alone, then transform the test set. Fitting on everything leaks test information into your model.
scaler.fit(X_train)
X_test_scaled = scaler.transform(X_test)Scale Inside a Pipeline
Bundle the scaler with your model in a Pipeline. It then fits only on the training fold during cross-validation, blocking leakage for free.
from sklearn.pipeline import make_pipeline
pipe = make_pipeline(StandardScaler(), model)Trees Do Not Need It
Tree-based models like random forests split on thresholds, so scaling rarely changes results. Save scaling for distance and linear methods.
Quick Check
You need every feature rescaled to a strict 0-to-1 range. Which scaler fits?
Recap: Scaling
You put features on a fair footing: StandardScaler for z-scores, MinMaxScaler for 0 to 1, RobustScaler for outliers. Always fit on train only. 🎉
الأسئلة الشائعة
هل درس «تحجيم الأرقام وتطبيعها» مجاني؟
نعم — نص درس «تحجيم الأرقام وتطبيعها» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Data Science Academy، انتقل إلى CoddyKit PRO. تتضمن دورة Data Science Academy 4 دروس في المجموع.
ماذا ستتعلم في «تحجيم الأرقام وتطبيعها»؟
وضع الميزات على مقياس عادل تتمرن على Data Science Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ Data Science Academy؟
لا تُشترط خبرة سابقة. Data Science Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 3 من أصل 4.
كم من الوقت يستغرق درس «تحجيم الأرقام وتطبيعها»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس Data Science Academy هذا؟
نعم. كل درس في Data Science Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- تحويل الأرقام إلى فئات
- ترميز الأعمدة الفئوية
- تحجيم الأرقام وتطبيعها
- إنشاء ميزات من التواريخ والنصوص