0Pricing
Machine Learning Academy · 课时

使用 Pandas 操作数据

您将把 CSV 文件加载到 DataFrames 中,筛选行、选择列、处理缺失值,并计算汇总统计量

使用 Pandas 操作数据 是 CoddyKit 上的免费 Machine Learning Academy 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Machine Learning Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Machine Learning Academy 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

What Is Pandas and Why Use It?

Pandas handles tabular data — rows and columns, like a spreadsheet. Its DataFrame is where you load and clean raw data before it ever reaches your model.

import pandas as pd
import numpy as np

# Create a DataFrame manually
df = pd.DataFrame({
    'name': ['Alice', 'Bob', 'Carol', 'Dave'],
    'age': [25, 30, 35, 28],
    'salary': [50000, 70000, 90000, 60000],
    'department': ['Engineering', 'Marketing', 'Engineering', 'Sales']
})
print(df)
print('\nShape:', df.shape)  # (4, 4)

Loading Data from CSV Files

Load a file in one line with pd.read_csv(). Then always inspect it: .head(), .info(), and .describe() reveal shape, types, and missing values in seconds.

import pandas as pd

# Load a CSV
df = pd.read_csv('titanic.csv')

# First inspection
print(df.head())       # first 5 rows
print(df.tail(3))      # last 3 rows
print(df.info())       # column types + non-null counts
print(df.describe())   # count, mean, std, min, max for numeric cols
print(df.columns.tolist())  # column names
print(df.shape)        # (891, 12)

Selecting Columns and Rows

Pick columns with df['col'], and filter rows with boolean conditions. Use .loc[] for labels and .iloc[] for positions. Wrap each condition in parentheses!

import pandas as pd

df = pd.read_csv('titanic.csv')

# Select columns
age = df['Age']                          # Series
subset = df[['Age', 'Fare', 'Survived']] # DataFrame

# Filter rows with boolean conditions
survivors = df[df['Survived'] == 1]
first_class_women = df[(df['Pclass'] == 1) & (df['Sex'] == 'female')]

# Label-based selection (row label, column label)
print(df.loc[0, 'Name'])  # first row, Name column

# Position-based selection
print(df.iloc[0, 0])      # first row, first column

Handling Missing Values

Real data has gaps, shown as NaN, and most models choke on them. Your two moves: drop the rows or columns, or fill them with a value like the median.

import pandas as pd
import numpy as np

df = pd.read_csv('titanic.csv')

# Detect missing values
print(df.isnull().sum())  # count NaN per column
print(df.isnull().mean() * 100)  # % missing per column

# Drop columns with >50% missing
df_clean = df.dropna(thresh=len(df) * 0.5, axis=1)

# Fill numeric missing with median
df['Age'] = df['Age'].fillna(df['Age'].median())

# Fill categorical missing with most frequent
df['Embarked'] = df['Embarked'].fillna(df['Embarked'].mode()[0])

Data Types and Type Conversion

Every column has a dtype like int, float, or object (text). A wrong one causes silent bugs — numbers read as text can't do math. Use .astype() to convert.

import pandas as pd

df = pd.read_csv('titanic.csv')
print(df.dtypes)  # see all column dtypes

# Convert a column's dtype
df['Survived'] = df['Survived'].astype(bool)
df['Pclass'] = df['Pclass'].astype('category')

# Convert object column to numeric (coerce errors to NaN)
df['Fare'] = pd.to_numeric(df['Fare'], errors='coerce')

# Parse dates
# df['date'] = pd.to_datetime(df['date'])

print(df.dtypes)

Adding and Transforming Columns

Feature engineering means building new columns from existing ones. Combining SibSp and Parch into one FamilySize often helps a model learn better. See the code.

import pandas as pd

df = pd.read_csv('titanic.csv')

# Create a new column from existing ones
df['FamilySize'] = df['SibSp'] + df['Parch'] + 1  # +1 for self

# Binary flag: is the passenger alone?
df['IsAlone'] = (df['FamilySize'] == 1).astype(int)

# Bin a continuous feature into categories
df['AgeGroup'] = pd.cut(df['Age'], bins=[0, 12, 18, 60, 100],
                         labels=['Child', 'Teen', 'Adult', 'Senior'])

print(df[['FamilySize', 'IsAlone', 'AgeGroup']].head())

GroupBy: Aggregating by Category

groupby() splits your data by a category, applies a function like mean, and combines the results — just like SQL's GROUP BY. Perfect for spotting patterns fast.

import pandas as pd

df = pd.read_csv('titanic.csv')

# Mean survival rate by class
survival_by_class = df.groupby('Pclass')['Survived'].mean()
print(survival_by_class)
# Pclass 1: ~0.63, Pclass 2: ~0.47, Pclass 3: ~0.24

# Multiple aggregations at once
summary = df.groupby('Pclass').agg(
    passengers=('Survived', 'count'),
    survival_rate=('Survived', 'mean'),
    avg_fare=('Fare', 'mean')
)
print(summary)

Merging and Joining DataFrames

Need data from two sources? pd.merge() joins DataFrames on a shared key, just like a SQL join. Use pd.concat() to stack tables into more rows or columns.

import pandas as pd

customers = pd.DataFrame({'id': [1, 2, 3], 'name': ['Alice', 'Bob', 'Carol']})
orders = pd.DataFrame({'id': [1, 1, 2], 'amount': [50, 30, 70]})

# Inner join on 'id'
merged = pd.merge(customers, orders, on='id', how='inner')
print(merged)

# Stack DataFrames with the same columns
df1 = pd.DataFrame({'A': [1, 2]})
df2 = pd.DataFrame({'A': [3, 4]})
combined = pd.concat([df1, df2], ignore_index=True)
print(combined)

Summary Statistics and Value Counts

Before modeling, study your data: .describe() summarizes numbers and .value_counts() counts categories. Always check for class imbalance — it can fool accuracy.

import pandas as pd

df = pd.read_csv('titanic.csv')

# Numeric statistics
print(df['Age'].describe())
# count, mean, std, min, 25%, 50%, 75%, max

# Categorical distribution
print(df['Sex'].value_counts())
print(df['Pclass'].value_counts(normalize=True))  # proportions

# Correlation with target variable
correlations = df.corr()['Survived'].sort_values(ascending=False)
print(correlations)

Converting a DataFrame to NumPy for scikit-learn

The last step before training: split features (X) from the target (y), then convert to NumPy with .to_numpy(). Now scikit-learn has exactly what it expects.

import pandas as pd
from sklearn.model_selection import train_test_split

df = pd.read_csv('titanic.csv')

# Select numeric features only for simplicity
features = ['Pclass', 'Age', 'SibSp', 'Parch', 'Fare']
df_model = df[features + ['Survived']].dropna()

X = df_model[features].to_numpy()       # shape (n, 5)
y = df_model['Survived'].to_numpy()     # shape (n,)

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)
print('Train shape:', X_train.shape)
print('Test shape:', X_test.shape)

Saving and Loading Processed Data

After all that cleaning, save your work so you never redo it. df.to_csv() is simple; Parquet is far smaller and faster for big datasets.

import pandas as pd

df = pd.read_csv('titanic.csv')

# Save as CSV
df.to_csv('titanic_processed.csv', index=False)

# Save as Parquet (much faster for large files)
# df.to_parquet('titanic_processed.parquet', index=False)

# Load back
df_reloaded = pd.read_csv('titanic_processed.csv')
print('Reloaded shape:', df_reloaded.shape)

Quick Check

Test your understanding of Machine Learning with Python concepts from this lesson.

Lesson Recap

You learned to wrangle data with Pandas: DataFrames load and clean tables, dropna and fillna handle missing values, and groupby reveals patterns. Next: charts. 📊

常见问题解答

「使用 Pandas 操作数据」课时是免费的吗?

是的 — 「使用 Pandas 操作数据」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Machine Learning Academy 课程的其余内容,请升级到 CoddyKit PRO。 Machine Learning Academy 课程共包含 4 节课。

「使用 Pandas 操作数据」这节课中我会学到什么?

您将把 CSV 文件加载到 DataFrames 中,筛选行、选择列、处理缺失值,并计算汇总统计量 你通过在浏览器中直接运行的动手代码来练习 Machine Learning Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Machine Learning Academy 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Machine Learning Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。

「使用 Pandas 操作数据」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Machine Learning Academy 课中编写并运行代码吗?

能。每节 Machine Learning Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 安装 Anaconda 与 Jupyter Notebook
  2. NumPy 基础:数组与数学运算
  3. 使用 Pandas 操作数据
  4. 使用 Matplotlib 和 Seaborn 可视化数据
← 返回 Machine Learning Academy