0Pricing
Machine Learning Academy · Lesson

Pandas for Data Manipulation

Learners will load CSV files into DataFrames, filter rows, select columns, handle missing values, and compute summary statistics.

Pandas for Data Manipulation is a free Machine Learning Academy lesson on CoddyKit — lesson 3 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Machine Learning Academy learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.

What Is Pandas and Why Use It?

Pandas handles tabular data — rows and columns, like a spreadsheet. Its DataFrame is where you load and clean raw data before it ever reaches your model.

import pandas as pd
import numpy as np

# Create a DataFrame manually
df = pd.DataFrame({
    'name': ['Alice', 'Bob', 'Carol', 'Dave'],
    'age': [25, 30, 35, 28],
    'salary': [50000, 70000, 90000, 60000],
    'department': ['Engineering', 'Marketing', 'Engineering', 'Sales']
})
print(df)
print('\nShape:', df.shape)  # (4, 4)

Loading Data from CSV Files

Load a file in one line with pd.read_csv(). Then always inspect it: .head(), .info(), and .describe() reveal shape, types, and missing values in seconds.

import pandas as pd

# Load a CSV
df = pd.read_csv('titanic.csv')

# First inspection
print(df.head())       # first 5 rows
print(df.tail(3))      # last 3 rows
print(df.info())       # column types + non-null counts
print(df.describe())   # count, mean, std, min, max for numeric cols
print(df.columns.tolist())  # column names
print(df.shape)        # (891, 12)

Selecting Columns and Rows

Pick columns with df['col'], and filter rows with boolean conditions. Use .loc[] for labels and .iloc[] for positions. Wrap each condition in parentheses!

import pandas as pd

df = pd.read_csv('titanic.csv')

# Select columns
age = df['Age']                          # Series
subset = df[['Age', 'Fare', 'Survived']] # DataFrame

# Filter rows with boolean conditions
survivors = df[df['Survived'] == 1]
first_class_women = df[(df['Pclass'] == 1) & (df['Sex'] == 'female')]

# Label-based selection (row label, column label)
print(df.loc[0, 'Name'])  # first row, Name column

# Position-based selection
print(df.iloc[0, 0])      # first row, first column

Handling Missing Values

Real data has gaps, shown as NaN, and most models choke on them. Your two moves: drop the rows or columns, or fill them with a value like the median.

import pandas as pd
import numpy as np

df = pd.read_csv('titanic.csv')

# Detect missing values
print(df.isnull().sum())  # count NaN per column
print(df.isnull().mean() * 100)  # % missing per column

# Drop columns with >50% missing
df_clean = df.dropna(thresh=len(df) * 0.5, axis=1)

# Fill numeric missing with median
df['Age'] = df['Age'].fillna(df['Age'].median())

# Fill categorical missing with most frequent
df['Embarked'] = df['Embarked'].fillna(df['Embarked'].mode()[0])

Data Types and Type Conversion

Every column has a dtype like int, float, or object (text). A wrong one causes silent bugs — numbers read as text can't do math. Use .astype() to convert.

import pandas as pd

df = pd.read_csv('titanic.csv')
print(df.dtypes)  # see all column dtypes

# Convert a column's dtype
df['Survived'] = df['Survived'].astype(bool)
df['Pclass'] = df['Pclass'].astype('category')

# Convert object column to numeric (coerce errors to NaN)
df['Fare'] = pd.to_numeric(df['Fare'], errors='coerce')

# Parse dates
# df['date'] = pd.to_datetime(df['date'])

print(df.dtypes)

Adding and Transforming Columns

Feature engineering means building new columns from existing ones. Combining SibSp and Parch into one FamilySize often helps a model learn better. See the code.

import pandas as pd

df = pd.read_csv('titanic.csv')

# Create a new column from existing ones
df['FamilySize'] = df['SibSp'] + df['Parch'] + 1  # +1 for self

# Binary flag: is the passenger alone?
df['IsAlone'] = (df['FamilySize'] == 1).astype(int)

# Bin a continuous feature into categories
df['AgeGroup'] = pd.cut(df['Age'], bins=[0, 12, 18, 60, 100],
                         labels=['Child', 'Teen', 'Adult', 'Senior'])

print(df[['FamilySize', 'IsAlone', 'AgeGroup']].head())

GroupBy: Aggregating by Category

groupby() splits your data by a category, applies a function like mean, and combines the results — just like SQL's GROUP BY. Perfect for spotting patterns fast.

import pandas as pd

df = pd.read_csv('titanic.csv')

# Mean survival rate by class
survival_by_class = df.groupby('Pclass')['Survived'].mean()
print(survival_by_class)
# Pclass 1: ~0.63, Pclass 2: ~0.47, Pclass 3: ~0.24

# Multiple aggregations at once
summary = df.groupby('Pclass').agg(
    passengers=('Survived', 'count'),
    survival_rate=('Survived', 'mean'),
    avg_fare=('Fare', 'mean')
)
print(summary)

Merging and Joining DataFrames

Need data from two sources? pd.merge() joins DataFrames on a shared key, just like a SQL join. Use pd.concat() to stack tables into more rows or columns.

import pandas as pd

customers = pd.DataFrame({'id': [1, 2, 3], 'name': ['Alice', 'Bob', 'Carol']})
orders = pd.DataFrame({'id': [1, 1, 2], 'amount': [50, 30, 70]})

# Inner join on 'id'
merged = pd.merge(customers, orders, on='id', how='inner')
print(merged)

# Stack DataFrames with the same columns
df1 = pd.DataFrame({'A': [1, 2]})
df2 = pd.DataFrame({'A': [3, 4]})
combined = pd.concat([df1, df2], ignore_index=True)
print(combined)

Summary Statistics and Value Counts

Before modeling, study your data: .describe() summarizes numbers and .value_counts() counts categories. Always check for class imbalance — it can fool accuracy.

import pandas as pd

df = pd.read_csv('titanic.csv')

# Numeric statistics
print(df['Age'].describe())
# count, mean, std, min, 25%, 50%, 75%, max

# Categorical distribution
print(df['Sex'].value_counts())
print(df['Pclass'].value_counts(normalize=True))  # proportions

# Correlation with target variable
correlations = df.corr()['Survived'].sort_values(ascending=False)
print(correlations)

Converting a DataFrame to NumPy for scikit-learn

The last step before training: split features (X) from the target (y), then convert to NumPy with .to_numpy(). Now scikit-learn has exactly what it expects.

import pandas as pd
from sklearn.model_selection import train_test_split

df = pd.read_csv('titanic.csv')

# Select numeric features only for simplicity
features = ['Pclass', 'Age', 'SibSp', 'Parch', 'Fare']
df_model = df[features + ['Survived']].dropna()

X = df_model[features].to_numpy()       # shape (n, 5)
y = df_model['Survived'].to_numpy()     # shape (n,)

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)
print('Train shape:', X_train.shape)
print('Test shape:', X_test.shape)

Saving and Loading Processed Data

After all that cleaning, save your work so you never redo it. df.to_csv() is simple; Parquet is far smaller and faster for big datasets.

import pandas as pd

df = pd.read_csv('titanic.csv')

# Save as CSV
df.to_csv('titanic_processed.csv', index=False)

# Save as Parquet (much faster for large files)
# df.to_parquet('titanic_processed.parquet', index=False)

# Load back
df_reloaded = pd.read_csv('titanic_processed.csv')
print('Reloaded shape:', df_reloaded.shape)

Quick Check

Test your understanding of Machine Learning with Python concepts from this lesson.

Lesson Recap

You learned to wrangle data with Pandas: DataFrames load and clean tables, dropna and fillna handle missing values, and groupby reveals patterns. Next: charts. 📊

Frequently asked questions

Is the “Pandas for Data Manipulation” lesson free?

Yes — the full text of “Pandas for Data Manipulation” is free to read here on the web, and the Machine Learning Academy course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Machine Learning Academy course, upgrade to CoddyKit PRO.

What will I learn in “Pandas for Data Manipulation”?

Learners will load CSV files into DataFrames, filter rows, select columns, handle missing values, and compute summary statistics. You practise Machine Learning Academy with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.

Do I need any experience to start Machine Learning Academy?

No prior experience is required. Machine Learning Academy on CoddyKit is structured for beginners through advanced learners; this is — lesson 3 of 4, so you can start here or from the beginning and move at your own pace.

How long does the “Pandas for Data Manipulation” lesson take?

Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.

Can I write and run code in this Machine Learning Academy lesson?

Yes. Every Machine Learning Academy lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.

All lessons in this course

  1. Installing Anaconda and Jupyter Notebook
  2. NumPy Essentials: Arrays and Math Operations
  3. Pandas for Data Manipulation
  4. Visualising Data with Matplotlib and Seaborn
← Back to Machine Learning Academy