Pandas for Data Manipulation
Learners will load CSV files into DataFrames, filter rows, select columns, handle missing values, and compute summary statistics.
Pandas for Data Manipulation is a free Machine Learning Academy lesson on CoddyKit — lesson 3 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Machine Learning Academy learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
What Is Pandas and Why Use It?
Pandas handles tabular data — rows and columns, like a spreadsheet. Its DataFrame is where you load and clean raw data before it ever reaches your model.
import pandas as pd
import numpy as np
# Create a DataFrame manually
df = pd.DataFrame({
'name': ['Alice', 'Bob', 'Carol', 'Dave'],
'age': [25, 30, 35, 28],
'salary': [50000, 70000, 90000, 60000],
'department': ['Engineering', 'Marketing', 'Engineering', 'Sales']
})
print(df)
print('\nShape:', df.shape) # (4, 4)Loading Data from CSV Files
Load a file in one line with pd.read_csv(). Then always inspect it: .head(), .info(), and .describe() reveal shape, types, and missing values in seconds.
import pandas as pd
# Load a CSV
df = pd.read_csv('titanic.csv')
# First inspection
print(df.head()) # first 5 rows
print(df.tail(3)) # last 3 rows
print(df.info()) # column types + non-null counts
print(df.describe()) # count, mean, std, min, max for numeric cols
print(df.columns.tolist()) # column names
print(df.shape) # (891, 12)Selecting Columns and Rows
Pick columns with df['col'], and filter rows with boolean conditions. Use .loc[] for labels and .iloc[] for positions. Wrap each condition in parentheses!
import pandas as pd
df = pd.read_csv('titanic.csv')
# Select columns
age = df['Age'] # Series
subset = df[['Age', 'Fare', 'Survived']] # DataFrame
# Filter rows with boolean conditions
survivors = df[df['Survived'] == 1]
first_class_women = df[(df['Pclass'] == 1) & (df['Sex'] == 'female')]
# Label-based selection (row label, column label)
print(df.loc[0, 'Name']) # first row, Name column
# Position-based selection
print(df.iloc[0, 0]) # first row, first columnHandling Missing Values
Real data has gaps, shown as NaN, and most models choke on them. Your two moves: drop the rows or columns, or fill them with a value like the median.
import pandas as pd
import numpy as np
df = pd.read_csv('titanic.csv')
# Detect missing values
print(df.isnull().sum()) # count NaN per column
print(df.isnull().mean() * 100) # % missing per column
# Drop columns with >50% missing
df_clean = df.dropna(thresh=len(df) * 0.5, axis=1)
# Fill numeric missing with median
df['Age'] = df['Age'].fillna(df['Age'].median())
# Fill categorical missing with most frequent
df['Embarked'] = df['Embarked'].fillna(df['Embarked'].mode()[0])Data Types and Type Conversion
Every column has a dtype like int, float, or object (text). A wrong one causes silent bugs — numbers read as text can't do math. Use .astype() to convert.
import pandas as pd
df = pd.read_csv('titanic.csv')
print(df.dtypes) # see all column dtypes
# Convert a column's dtype
df['Survived'] = df['Survived'].astype(bool)
df['Pclass'] = df['Pclass'].astype('category')
# Convert object column to numeric (coerce errors to NaN)
df['Fare'] = pd.to_numeric(df['Fare'], errors='coerce')
# Parse dates
# df['date'] = pd.to_datetime(df['date'])
print(df.dtypes)Adding and Transforming Columns
Feature engineering means building new columns from existing ones. Combining SibSp and Parch into one FamilySize often helps a model learn better. See the code.
import pandas as pd
df = pd.read_csv('titanic.csv')
# Create a new column from existing ones
df['FamilySize'] = df['SibSp'] + df['Parch'] + 1 # +1 for self
# Binary flag: is the passenger alone?
df['IsAlone'] = (df['FamilySize'] == 1).astype(int)
# Bin a continuous feature into categories
df['AgeGroup'] = pd.cut(df['Age'], bins=[0, 12, 18, 60, 100],
labels=['Child', 'Teen', 'Adult', 'Senior'])
print(df[['FamilySize', 'IsAlone', 'AgeGroup']].head())GroupBy: Aggregating by Category
groupby() splits your data by a category, applies a function like mean, and combines the results — just like SQL's GROUP BY. Perfect for spotting patterns fast.
import pandas as pd
df = pd.read_csv('titanic.csv')
# Mean survival rate by class
survival_by_class = df.groupby('Pclass')['Survived'].mean()
print(survival_by_class)
# Pclass 1: ~0.63, Pclass 2: ~0.47, Pclass 3: ~0.24
# Multiple aggregations at once
summary = df.groupby('Pclass').agg(
passengers=('Survived', 'count'),
survival_rate=('Survived', 'mean'),
avg_fare=('Fare', 'mean')
)
print(summary)Merging and Joining DataFrames
Need data from two sources? pd.merge() joins DataFrames on a shared key, just like a SQL join. Use pd.concat() to stack tables into more rows or columns.
import pandas as pd
customers = pd.DataFrame({'id': [1, 2, 3], 'name': ['Alice', 'Bob', 'Carol']})
orders = pd.DataFrame({'id': [1, 1, 2], 'amount': [50, 30, 70]})
# Inner join on 'id'
merged = pd.merge(customers, orders, on='id', how='inner')
print(merged)
# Stack DataFrames with the same columns
df1 = pd.DataFrame({'A': [1, 2]})
df2 = pd.DataFrame({'A': [3, 4]})
combined = pd.concat([df1, df2], ignore_index=True)
print(combined)Summary Statistics and Value Counts
Before modeling, study your data: .describe() summarizes numbers and .value_counts() counts categories. Always check for class imbalance — it can fool accuracy.
import pandas as pd
df = pd.read_csv('titanic.csv')
# Numeric statistics
print(df['Age'].describe())
# count, mean, std, min, 25%, 50%, 75%, max
# Categorical distribution
print(df['Sex'].value_counts())
print(df['Pclass'].value_counts(normalize=True)) # proportions
# Correlation with target variable
correlations = df.corr()['Survived'].sort_values(ascending=False)
print(correlations)Converting a DataFrame to NumPy for scikit-learn
The last step before training: split features (X) from the target (y), then convert to NumPy with .to_numpy(). Now scikit-learn has exactly what it expects.
import pandas as pd
from sklearn.model_selection import train_test_split
df = pd.read_csv('titanic.csv')
# Select numeric features only for simplicity
features = ['Pclass', 'Age', 'SibSp', 'Parch', 'Fare']
df_model = df[features + ['Survived']].dropna()
X = df_model[features].to_numpy() # shape (n, 5)
y = df_model['Survived'].to_numpy() # shape (n,)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
print('Train shape:', X_train.shape)
print('Test shape:', X_test.shape)Saving and Loading Processed Data
After all that cleaning, save your work so you never redo it. df.to_csv() is simple; Parquet is far smaller and faster for big datasets.
import pandas as pd
df = pd.read_csv('titanic.csv')
# Save as CSV
df.to_csv('titanic_processed.csv', index=False)
# Save as Parquet (much faster for large files)
# df.to_parquet('titanic_processed.parquet', index=False)
# Load back
df_reloaded = pd.read_csv('titanic_processed.csv')
print('Reloaded shape:', df_reloaded.shape)Quick Check
Test your understanding of Machine Learning with Python concepts from this lesson.
Lesson Recap
You learned to wrangle data with Pandas: DataFrames load and clean tables, dropna and fillna handle missing values, and groupby reveals patterns. Next: charts. 📊
Frequently asked questions
Is the “Pandas for Data Manipulation” lesson free?
Yes — the full text of “Pandas for Data Manipulation” is free to read here on the web, and the Machine Learning Academy course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Machine Learning Academy course, upgrade to CoddyKit PRO.
What will I learn in “Pandas for Data Manipulation”?
Learners will load CSV files into DataFrames, filter rows, select columns, handle missing values, and compute summary statistics. You practise Machine Learning Academy with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start Machine Learning Academy?
No prior experience is required. Machine Learning Academy on CoddyKit is structured for beginners through advanced learners; this is — lesson 3 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “Pandas for Data Manipulation” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this Machine Learning Academy lesson?
Yes. Every Machine Learning Academy lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.