使用 Matplotlib 和 Seaborn 可视化数据
您将绘制直方图、散点图和相关性热力图,在建模前探索数据分布与关系
使用 Matplotlib 和 Seaborn 可视化数据 是 CoddyKit 上的免费 Machine Learning Academy 课时。 这是第 4 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Machine Learning Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Machine Learning Academy 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Why Visualisation Matters in ML
Charts aren't just for reports — they're a diagnostic tool at every step. Matplotlib gives full control; Seaborn makes beautiful stats plots with less code.
import matplotlib.pyplot as plt
import seaborn as sns
import pandas as pd
import numpy as np
# Set seaborn theme for all plots
sns.set_theme(style='whitegrid', palette='muted')
# Load a built-in dataset
df = sns.load_dataset('tips')
print(df.head())Histograms: Understanding Distributions
A histogram bins a numeric column to show its shape — normal, skewed, or with outliers. Spotting skew tells you when a log transform might help your model.
import matplotlib.pyplot as plt
import seaborn as sns
df = sns.load_dataset('tips')
fig, axes = plt.subplots(1, 2, figsize=(12, 4))
# Raw distribution (right-skewed)
axes[0].hist(df['total_bill'], bins=30, edgecolor='white')
axes[0].set_title('Total Bill Distribution (Raw)')
# After log transform
import numpy as np
axes[1].hist(np.log(df['total_bill']), bins=30, edgecolor='white')
axes[1].set_title('Total Bill Distribution (Log Transformed)')
plt.tight_layout()
plt.show()Scatter Plots: Feature Relationships
A scatter plot shows how two numbers relate — linear, curved, or full of outliers. Add color with hue to squeeze a third variable into the same view.
import matplotlib.pyplot as plt
import seaborn as sns
df = sns.load_dataset('tips')
# Scatter plot with hue encoding
sns.scatterplot(
data=df,
x='total_bill',
y='tip',
hue='time', # encode meal time as colour
size='size', # encode party size as dot size
alpha=0.7
)
plt.title('Tip vs Total Bill (coloured by Meal Time)')
plt.show()Box Plots: Comparing Groups
A box plot shows the median, spread, and outliers across groups. If a feature's box looks very different per category, that feature is probably worth keeping.
import matplotlib.pyplot as plt
import seaborn as sns
df = sns.load_dataset('tips')
fig, axes = plt.subplots(1, 2, figsize=(12, 5))
sns.boxplot(data=df, x='day', y='total_bill', ax=axes[0])
axes[0].set_title('Bill by Day of Week')
sns.boxplot(data=df, x='smoker', y='tip', hue='sex', ax=axes[1])
axes[1].set_title('Tip by Smoking Status and Sex')
plt.tight_layout()
plt.show()Correlation Heatmap: Finding Related Features
A correlation heatmap colors how strongly every pair of columns moves together. It reveals good predictors of your target — and redundant features to drop.
import matplotlib.pyplot as plt
import seaborn as sns
import pandas as pd
df = pd.read_csv('titanic.csv')
# Compute correlation matrix
corr = df[['Survived', 'Pclass', 'Age', 'SibSp', 'Parch', 'Fare']].corr()
# Plot heatmap
plt.figure(figsize=(8, 6))
sns.heatmap(
corr,
annot=True, # show correlation values
fmt='.2f', # 2 decimal places
cmap='RdYlGn', # red-yellow-green colour scale
vmin=-1, vmax=1
)
plt.title('Feature Correlation Matrix')
plt.show()Pair Plots: Exploring All Feature Pairs
A pair plot shows scatter plots for every feature pair at once — a fast first look at a new dataset. Color by class to see which features separate the groups.
import seaborn as sns
import matplotlib.pyplot as plt
# Use the iris dataset (classic ML benchmark)
df = sns.load_dataset('iris')
# Pair plot coloured by species
sns.pairplot(
df,
hue='species',
diag_kind='kde', # KDE on diagonal
plot_kws={'alpha': 0.6}
)
plt.suptitle('Iris Dataset Pair Plot', y=1.02)
plt.show()Bar Charts and Count Plots
A count plot shows how often each category appears — the quickest way to check class balance before training a classifier. Add hue to compare two categories.
import seaborn as sns
import matplotlib.pyplot as plt
df = sns.load_dataset('titanic')
fig, axes = plt.subplots(1, 2, figsize=(12, 5))
# Class distribution (check balance)
sns.countplot(data=df, x='survived', ax=axes[0])
axes[0].set_title('Survival Count')
# Class by passenger class and sex
sns.countplot(data=df, x='class', hue='sex', ax=axes[1])
axes[1].set_title('Class Distribution by Sex')
plt.tight_layout()
plt.show()Plotting Learning Curves
A learning curve plots training vs validation score as data grows. Both low means underfitting; a big gap means overfitting; both high and close means a good fit.
import matplotlib.pyplot as plt
import numpy as np
from sklearn.model_selection import learning_curve
from sklearn.tree import DecisionTreeClassifier
from sklearn.datasets import load_breast_cancer
X, y = load_breast_cancer(return_X_y=True)
train_sizes, train_scores, val_scores = learning_curve(
DecisionTreeClassifier(max_depth=5), X, y, cv=5
)
plt.plot(train_sizes, train_scores.mean(axis=1), label='Training Score')
plt.plot(train_sizes, val_scores.mean(axis=1), label='Validation Score')
plt.xlabel('Training Set Size')
plt.ylabel('Accuracy')
plt.legend()
plt.title('Learning Curve')
plt.show()Visualising Model Predictions
After training, plot the predictions. A confusion matrix heatmap shows where a classifier confuses labels — far more telling than a single accuracy number.
import seaborn as sns
import matplotlib.pyplot as plt
from sklearn.metrics import confusion_matrix
from sklearn.datasets import load_iris
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
model = DecisionTreeClassifier(max_depth=3)
model.fit(X_train, y_train)
y_pred = model.predict(X_test)
cm = confusion_matrix(y_test, y_pred)
sns.heatmap(cm, annot=True, fmt='d', cmap='Blues')
plt.xlabel('Predicted')
plt.ylabel('Actual')
plt.title('Confusion Matrix')
plt.show()Saving and Customising Plots
Good plots need polish: titles, axis labels, and readable fonts. Save them with savefig at dpi=150+ for reports, and pick a colorblind-friendly palette.
import matplotlib.pyplot as plt
import seaborn as sns
import numpy as np
# Create figure and axes explicitly
fig, ax = plt.subplots(figsize=(8, 5))
x = np.linspace(0, 10, 100)
ax.plot(x, np.sin(x), label='sin(x)', linewidth=2)
ax.plot(x, np.cos(x), label='cos(x)', linewidth=2, linestyle='--')
# Customise
ax.set_xlabel('x', fontsize=13)
ax.set_ylabel('y', fontsize=13)
ax.set_title('Sine and Cosine', fontsize=15, fontweight='bold')
ax.legend(fontsize=12)
ax.grid(True, alpha=0.3)
# Save
fig.savefig('plot.png', dpi=150, bbox_inches='tight')
plt.show()Distribution Plots with Seaborn
A violin plot blends a box plot with a density curve, showing the full shape of a distribution. Seaborn's displot and kdeplot are great for smooth comparisons.
import seaborn as sns
import matplotlib.pyplot as plt
df = sns.load_dataset('tips')
fig, axes = plt.subplots(1, 2, figsize=(12, 5))
# Violin plot
sns.violinplot(data=df, x='day', y='total_bill', hue='sex',
split=True, inner='quart', ax=axes[0])
axes[0].set_title('Bill Distribution by Day (Violin)')
# KDE distribution comparison
sns.kdeplot(data=df, x='tip', hue='time', fill=True, alpha=0.4, ax=axes[1])
axes[1].set_title('Tip Distribution by Meal Time (KDE)')
plt.tight_layout()
plt.show()Quick Check
Test your understanding of Machine Learning with Python concepts from this lesson.
Lesson Recap
You learned to see your data: histograms and scatter plots reveal shape, heatmaps find predictors, and learning curves diagnose fit. Next: your first model! 🚀
常见问题解答
「使用 Matplotlib 和 Seaborn 可视化数据」课时是免费的吗?
是的 — 「使用 Matplotlib 和 Seaborn 可视化数据」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Machine Learning Academy 课程的其余内容,请升级到 CoddyKit PRO。 Machine Learning Academy 课程共包含 4 节课。
「使用 Matplotlib 和 Seaborn 可视化数据」这节课中我会学到什么?
您将绘制直方图、散点图和相关性热力图,在建模前探索数据分布与关系 你通过在浏览器中直接运行的动手代码来练习 Machine Learning Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Machine Learning Academy 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Machine Learning Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 4 节课,共 4 节。
「使用 Matplotlib 和 Seaborn 可视化数据」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Machine Learning Academy 课中编写并运行代码吗?
能。每节 Machine Learning Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- 安装 Anaconda 与 Jupyter Notebook
- NumPy 基础:数组与数学运算
- 使用 Pandas 操作数据
- 使用 Matplotlib 和 Seaborn 可视化数据