0Pricing
Machine Learning Academy · درس

استخدام Pandas لمعالجة البيانات

حمّل ملفات CSV إلى DataFrames، وصفِّ الصفوف، وحدد الأعمدة، وتعامل مع القيم المفقودة، واحسب الإحصاءات التلخيصية

استخدام Pandas لمعالجة البيانات درس مجاني في Machine Learning Academy على CoddyKit. هذا هو الدرس 3 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Machine Learning Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Machine Learning Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

What Is Pandas and Why Use It?

Pandas handles tabular data — rows and columns, like a spreadsheet. Its DataFrame is where you load and clean raw data before it ever reaches your model.

import pandas as pd
import numpy as np

# Create a DataFrame manually
df = pd.DataFrame({
    'name': ['Alice', 'Bob', 'Carol', 'Dave'],
    'age': [25, 30, 35, 28],
    'salary': [50000, 70000, 90000, 60000],
    'department': ['Engineering', 'Marketing', 'Engineering', 'Sales']
})
print(df)
print('\nShape:', df.shape)  # (4, 4)

Loading Data from CSV Files

Load a file in one line with pd.read_csv(). Then always inspect it: .head(), .info(), and .describe() reveal shape, types, and missing values in seconds.

import pandas as pd

# Load a CSV
df = pd.read_csv('titanic.csv')

# First inspection
print(df.head())       # first 5 rows
print(df.tail(3))      # last 3 rows
print(df.info())       # column types + non-null counts
print(df.describe())   # count, mean, std, min, max for numeric cols
print(df.columns.tolist())  # column names
print(df.shape)        # (891, 12)

Selecting Columns and Rows

Pick columns with df['col'], and filter rows with boolean conditions. Use .loc[] for labels and .iloc[] for positions. Wrap each condition in parentheses!

import pandas as pd

df = pd.read_csv('titanic.csv')

# Select columns
age = df['Age']                          # Series
subset = df[['Age', 'Fare', 'Survived']] # DataFrame

# Filter rows with boolean conditions
survivors = df[df['Survived'] == 1]
first_class_women = df[(df['Pclass'] == 1) & (df['Sex'] == 'female')]

# Label-based selection (row label, column label)
print(df.loc[0, 'Name'])  # first row, Name column

# Position-based selection
print(df.iloc[0, 0])      # first row, first column

Handling Missing Values

Real data has gaps, shown as NaN, and most models choke on them. Your two moves: drop the rows or columns, or fill them with a value like the median.

import pandas as pd
import numpy as np

df = pd.read_csv('titanic.csv')

# Detect missing values
print(df.isnull().sum())  # count NaN per column
print(df.isnull().mean() * 100)  # % missing per column

# Drop columns with >50% missing
df_clean = df.dropna(thresh=len(df) * 0.5, axis=1)

# Fill numeric missing with median
df['Age'] = df['Age'].fillna(df['Age'].median())

# Fill categorical missing with most frequent
df['Embarked'] = df['Embarked'].fillna(df['Embarked'].mode()[0])

Data Types and Type Conversion

Every column has a dtype like int, float, or object (text). A wrong one causes silent bugs — numbers read as text can't do math. Use .astype() to convert.

import pandas as pd

df = pd.read_csv('titanic.csv')
print(df.dtypes)  # see all column dtypes

# Convert a column's dtype
df['Survived'] = df['Survived'].astype(bool)
df['Pclass'] = df['Pclass'].astype('category')

# Convert object column to numeric (coerce errors to NaN)
df['Fare'] = pd.to_numeric(df['Fare'], errors='coerce')

# Parse dates
# df['date'] = pd.to_datetime(df['date'])

print(df.dtypes)

Adding and Transforming Columns

Feature engineering means building new columns from existing ones. Combining SibSp and Parch into one FamilySize often helps a model learn better. See the code.

import pandas as pd

df = pd.read_csv('titanic.csv')

# Create a new column from existing ones
df['FamilySize'] = df['SibSp'] + df['Parch'] + 1  # +1 for self

# Binary flag: is the passenger alone?
df['IsAlone'] = (df['FamilySize'] == 1).astype(int)

# Bin a continuous feature into categories
df['AgeGroup'] = pd.cut(df['Age'], bins=[0, 12, 18, 60, 100],
                         labels=['Child', 'Teen', 'Adult', 'Senior'])

print(df[['FamilySize', 'IsAlone', 'AgeGroup']].head())

GroupBy: Aggregating by Category

groupby() splits your data by a category, applies a function like mean, and combines the results — just like SQL's GROUP BY. Perfect for spotting patterns fast.

import pandas as pd

df = pd.read_csv('titanic.csv')

# Mean survival rate by class
survival_by_class = df.groupby('Pclass')['Survived'].mean()
print(survival_by_class)
# Pclass 1: ~0.63, Pclass 2: ~0.47, Pclass 3: ~0.24

# Multiple aggregations at once
summary = df.groupby('Pclass').agg(
    passengers=('Survived', 'count'),
    survival_rate=('Survived', 'mean'),
    avg_fare=('Fare', 'mean')
)
print(summary)

Merging and Joining DataFrames

Need data from two sources? pd.merge() joins DataFrames on a shared key, just like a SQL join. Use pd.concat() to stack tables into more rows or columns.

import pandas as pd

customers = pd.DataFrame({'id': [1, 2, 3], 'name': ['Alice', 'Bob', 'Carol']})
orders = pd.DataFrame({'id': [1, 1, 2], 'amount': [50, 30, 70]})

# Inner join on 'id'
merged = pd.merge(customers, orders, on='id', how='inner')
print(merged)

# Stack DataFrames with the same columns
df1 = pd.DataFrame({'A': [1, 2]})
df2 = pd.DataFrame({'A': [3, 4]})
combined = pd.concat([df1, df2], ignore_index=True)
print(combined)

Summary Statistics and Value Counts

Before modeling, study your data: .describe() summarizes numbers and .value_counts() counts categories. Always check for class imbalance — it can fool accuracy.

import pandas as pd

df = pd.read_csv('titanic.csv')

# Numeric statistics
print(df['Age'].describe())
# count, mean, std, min, 25%, 50%, 75%, max

# Categorical distribution
print(df['Sex'].value_counts())
print(df['Pclass'].value_counts(normalize=True))  # proportions

# Correlation with target variable
correlations = df.corr()['Survived'].sort_values(ascending=False)
print(correlations)

Converting a DataFrame to NumPy for scikit-learn

The last step before training: split features (X) from the target (y), then convert to NumPy with .to_numpy(). Now scikit-learn has exactly what it expects.

import pandas as pd
from sklearn.model_selection import train_test_split

df = pd.read_csv('titanic.csv')

# Select numeric features only for simplicity
features = ['Pclass', 'Age', 'SibSp', 'Parch', 'Fare']
df_model = df[features + ['Survived']].dropna()

X = df_model[features].to_numpy()       # shape (n, 5)
y = df_model['Survived'].to_numpy()     # shape (n,)

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)
print('Train shape:', X_train.shape)
print('Test shape:', X_test.shape)

Saving and Loading Processed Data

After all that cleaning, save your work so you never redo it. df.to_csv() is simple; Parquet is far smaller and faster for big datasets.

import pandas as pd

df = pd.read_csv('titanic.csv')

# Save as CSV
df.to_csv('titanic_processed.csv', index=False)

# Save as Parquet (much faster for large files)
# df.to_parquet('titanic_processed.parquet', index=False)

# Load back
df_reloaded = pd.read_csv('titanic_processed.csv')
print('Reloaded shape:', df_reloaded.shape)

Quick Check

Test your understanding of Machine Learning with Python concepts from this lesson.

Lesson Recap

You learned to wrangle data with Pandas: DataFrames load and clean tables, dropna and fillna handle missing values, and groupby reveals patterns. Next: charts. 📊

الأسئلة الشائعة

هل درس «استخدام Pandas لمعالجة البيانات» مجاني؟

نعم — نص درس «استخدام Pandas لمعالجة البيانات» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Machine Learning Academy، انتقل إلى CoddyKit PRO. تتضمن دورة Machine Learning Academy 4 دروس في المجموع.

ماذا ستتعلم في «استخدام Pandas لمعالجة البيانات»؟

حمّل ملفات CSV إلى DataFrames، وصفِّ الصفوف، وحدد الأعمدة، وتعامل مع القيم المفقودة، واحسب الإحصاءات التلخيصية تتمرن على Machine Learning Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ Machine Learning Academy؟

لا تُشترط خبرة سابقة. Machine Learning Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 3 من أصل 4.

كم من الوقت يستغرق درس «استخدام Pandas لمعالجة البيانات»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس Machine Learning Academy هذا؟

نعم. كل درس في Machine Learning Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. تثبيت Anaconda وJupyter Notebook
  2. أساسيات NumPy: المصفوفات والعمليات الحسابية
  3. استخدام Pandas لمعالجة البيانات
  4. تصور البيانات باستخدام Matplotlib وSeaborn
← العودة إلى Machine Learning Academy