Machine Learning Academy · Урок

Визуализация данных с Matplotlib и Seaborn

Вы построите гистограммы, диаграммы рассеяния и тепловые карты корреляций, чтобы изучить распределения и взаимосвязи данных перед моделированием

Урок 4 из 413 шагов

«Визуализация данных с Matplotlib и Seaborn» — бесплатный урок Machine Learning Academy на CoddyKit. Это урок 4 из 4. Ты можешь прочитать весь урок бесплатно ниже — а потом практиковать его прямо в браузере с встроенным редактором кода и ИИ-репетитором 24/7. Это часть пути обучения Machine Learning Academy, и твой прогресс синхронизируется между веб-версией и приложением CoddyKit. Курс Machine Learning Academy содержит 4 уроков всего.

Части этого урока еще не переведены и отображаются на английском.

Why Visualisation Matters in ML

Charts aren't just for reports — they're a diagnostic tool at every step. Matplotlib gives full control; Seaborn makes beautiful stats plots with less code.

import matplotlib.pyplot as plt
import seaborn as sns
import pandas as pd
import numpy as np

# Set seaborn theme for all plots
sns.set_theme(style='whitegrid', palette='muted')

# Load a built-in dataset
df = sns.load_dataset('tips')
print(df.head())

Histograms: Understanding Distributions

A histogram bins a numeric column to show its shape — normal, skewed, or with outliers. Spotting skew tells you when a log transform might help your model.

import matplotlib.pyplot as plt
import seaborn as sns

df = sns.load_dataset('tips')

fig, axes = plt.subplots(1, 2, figsize=(12, 4))

# Raw distribution (right-skewed)
axes[0].hist(df['total_bill'], bins=30, edgecolor='white')
axes[0].set_title('Total Bill Distribution (Raw)')

# After log transform
import numpy as np
axes[1].hist(np.log(df['total_bill']), bins=30, edgecolor='white')
axes[1].set_title('Total Bill Distribution (Log Transformed)')

plt.tight_layout()
plt.show()

Scatter Plots: Feature Relationships

A scatter plot shows how two numbers relate — linear, curved, or full of outliers. Add color with hue to squeeze a third variable into the same view.

import matplotlib.pyplot as plt
import seaborn as sns

df = sns.load_dataset('tips')

# Scatter plot with hue encoding
sns.scatterplot(
    data=df,
    x='total_bill',
    y='tip',
    hue='time',      # encode meal time as colour
    size='size',     # encode party size as dot size
    alpha=0.7
)
plt.title('Tip vs Total Bill (coloured by Meal Time)')
plt.show()

Box Plots: Comparing Groups

A box plot shows the median, spread, and outliers across groups. If a feature's box looks very different per category, that feature is probably worth keeping.

import matplotlib.pyplot as plt
import seaborn as sns

df = sns.load_dataset('tips')

fig, axes = plt.subplots(1, 2, figsize=(12, 5))

sns.boxplot(data=df, x='day', y='total_bill', ax=axes[0])
axes[0].set_title('Bill by Day of Week')

sns.boxplot(data=df, x='smoker', y='tip', hue='sex', ax=axes[1])
axes[1].set_title('Tip by Smoking Status and Sex')

plt.tight_layout()
plt.show()

Correlation Heatmap: Finding Related Features

A correlation heatmap colors how strongly every pair of columns moves together. It reveals good predictors of your target — and redundant features to drop.

import matplotlib.pyplot as plt
import seaborn as sns
import pandas as pd

df = pd.read_csv('titanic.csv')

# Compute correlation matrix
corr = df[['Survived', 'Pclass', 'Age', 'SibSp', 'Parch', 'Fare']].corr()

# Plot heatmap
plt.figure(figsize=(8, 6))
sns.heatmap(
    corr,
    annot=True,     # show correlation values
    fmt='.2f',      # 2 decimal places
    cmap='RdYlGn',  # red-yellow-green colour scale
    vmin=-1, vmax=1
)
plt.title('Feature Correlation Matrix')
plt.show()

Pair Plots: Exploring All Feature Pairs

A pair plot shows scatter plots for every feature pair at once — a fast first look at a new dataset. Color by class to see which features separate the groups.

import seaborn as sns
import matplotlib.pyplot as plt

# Use the iris dataset (classic ML benchmark)
df = sns.load_dataset('iris')

# Pair plot coloured by species
sns.pairplot(
    df,
    hue='species',
    diag_kind='kde',    # KDE on diagonal
    plot_kws={'alpha': 0.6}
)
plt.suptitle('Iris Dataset Pair Plot', y=1.02)
plt.show()

Bar Charts and Count Plots

A count plot shows how often each category appears — the quickest way to check class balance before training a classifier. Add hue to compare two categories.

import seaborn as sns
import matplotlib.pyplot as plt

df = sns.load_dataset('titanic')

fig, axes = plt.subplots(1, 2, figsize=(12, 5))

# Class distribution (check balance)
sns.countplot(data=df, x='survived', ax=axes[0])
axes[0].set_title('Survival Count')

# Class by passenger class and sex
sns.countplot(data=df, x='class', hue='sex', ax=axes[1])
axes[1].set_title('Class Distribution by Sex')

plt.tight_layout()
plt.show()

Plotting Learning Curves

A learning curve plots training vs validation score as data grows. Both low means underfitting; a big gap means overfitting; both high and close means a good fit.

import matplotlib.pyplot as plt
import numpy as np
from sklearn.model_selection import learning_curve
from sklearn.tree import DecisionTreeClassifier
from sklearn.datasets import load_breast_cancer

X, y = load_breast_cancer(return_X_y=True)
train_sizes, train_scores, val_scores = learning_curve(
    DecisionTreeClassifier(max_depth=5), X, y, cv=5
)

plt.plot(train_sizes, train_scores.mean(axis=1), label='Training Score')
plt.plot(train_sizes, val_scores.mean(axis=1), label='Validation Score')
plt.xlabel('Training Set Size')
plt.ylabel('Accuracy')
plt.legend()
plt.title('Learning Curve')
plt.show()

Visualising Model Predictions

After training, plot the predictions. A confusion matrix heatmap shows where a classifier confuses labels — far more telling than a single accuracy number.

import seaborn as sns
import matplotlib.pyplot as plt
from sklearn.metrics import confusion_matrix
from sklearn.datasets import load_iris
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split

X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)

model = DecisionTreeClassifier(max_depth=3)
model.fit(X_train, y_train)
y_pred = model.predict(X_test)

cm = confusion_matrix(y_test, y_pred)
sns.heatmap(cm, annot=True, fmt='d', cmap='Blues')
plt.xlabel('Predicted')
plt.ylabel('Actual')
plt.title('Confusion Matrix')
plt.show()

Saving and Customising Plots

Good plots need polish: titles, axis labels, and readable fonts. Save them with savefig at dpi=150+ for reports, and pick a colorblind-friendly palette.

import matplotlib.pyplot as plt
import seaborn as sns
import numpy as np

# Create figure and axes explicitly
fig, ax = plt.subplots(figsize=(8, 5))

x = np.linspace(0, 10, 100)
ax.plot(x, np.sin(x), label='sin(x)', linewidth=2)
ax.plot(x, np.cos(x), label='cos(x)', linewidth=2, linestyle='--')

# Customise
ax.set_xlabel('x', fontsize=13)
ax.set_ylabel('y', fontsize=13)
ax.set_title('Sine and Cosine', fontsize=15, fontweight='bold')
ax.legend(fontsize=12)
ax.grid(True, alpha=0.3)

# Save
fig.savefig('plot.png', dpi=150, bbox_inches='tight')
plt.show()

Distribution Plots with Seaborn

A violin plot blends a box plot with a density curve, showing the full shape of a distribution. Seaborn's displot and kdeplot are great for smooth comparisons.

import seaborn as sns
import matplotlib.pyplot as plt

df = sns.load_dataset('tips')

fig, axes = plt.subplots(1, 2, figsize=(12, 5))

# Violin plot
sns.violinplot(data=df, x='day', y='total_bill', hue='sex',
               split=True, inner='quart', ax=axes[0])
axes[0].set_title('Bill Distribution by Day (Violin)')

# KDE distribution comparison
sns.kdeplot(data=df, x='tip', hue='time', fill=True, alpha=0.4, ax=axes[1])
axes[1].set_title('Tip Distribution by Meal Time (KDE)')

plt.tight_layout()
plt.show()

Quick Check

Test your understanding of Machine Learning with Python concepts from this lesson.

Lesson Recap

You learned to see your data: histograms and scatter plots reveal shape, heatmaps find predictors, and learning curves diagnose fit. Next: your first model! 🚀

Можно начать бесплатно

Изучай Python с ИИ-репетитором — бесплатно

Пиши и запускай код прямо в браузере, получай мгновенную помощь от ИИ-репетитора 24/7 и продолжи учиться на сайте или в приложении.

Курсы
30
Уроки
120

Часто задаваемые вопросы

Урок «Визуализация данных с Matplotlib и Seaborn» бесплатный?

Да — полный текст урока «Визуализация данных с Matplotlib и Seaborn» бесплатно доступен здесь в веб-версии. Чтобы практиковать его интерактивно (встроенный редактор кода и ИИ-репетитор 24/7) и разблокировать остальной курс Machine Learning Academy, подпишись на CoddyKit PRO. Курс Machine Learning Academy содержит 4 уроков всего.

Чему я научусь в уроке «Визуализация данных с Matplotlib и Seaborn»?

Вы построите гистограммы, диаграммы рассеяния и тепловые карты корреляций, чтобы изучить распределения и взаимосвязи данных перед моделированием Ты практикуешь Machine Learning Academy с помощью реального кода, который запускаешь прямо в браузере, и ИИ-репетитор 24/7 отвечает на твои вопросы во время урока.

Нужен ли мне опыт, чтобы начать Machine Learning Academy?

Предыдущий опыт не требуется. Machine Learning Academy на CoddyKit структурирован для всех уровней — от новичков до продвинутых, поэтому ты можешь начать отсюда или с самого начала и учиться в своем темпе. Это урок 4 из 4.

Сколько времени занимает урок «Визуализация данных с Matplotlib и Seaborn»?

Большинство уроков CoddyKit занимают около 5–10 минут. Каждый из них компактный и интерактивный, поэтому ты постоянно делаешь прогресс и продолжаешь с того же места в веб-версии и приложении.

Можно ли писать и запускать код в этом уроке Machine Learning Academy?

Да. Каждый урок Machine Learning Academy включает встроенный редактор кода, поэтому ты пишешь и запускаешь реальный код прямо в браузере и получаешь моментальную обратную связь от AI — локальная установка не требуется.

Все уроки этого курса

  1. Установка Anaconda и Jupyter Notebook
  2. Основы NumPy: массивы и математические операции
  3. Pandas для обработки данных
  4. Визуализация данных с Matplotlib и Seaborn
← Назад к Machine Learning Academy