0Pricing
Pandas & NumPy Academy · Урок

Диаграммы рассеяния и парные графики

Создавайте диаграммы рассеяния, окрашенные по третьей переменной, с помощью sns.scatterplot и визуализируйте все попарные взаимосвязи с помощью sns.pairplot.

«Диаграммы рассеяния и парные графики» — бесплатный урок Pandas & NumPy Academy на CoddyKit. Это урок 3 из 4. Ты можешь прочитать весь урок бесплатно ниже — а потом практиковать его прямо в браузере с встроенным редактором кода и ИИ-репетитором 24/7. Это часть пути обучения Pandas & NumPy Academy, и твой прогресс синхронизируется между веб-версией и приложением CoddyKit. Курс Pandas & NumPy Academy содержит 4 уроков всего.

Части этого урока еще не переведены и отображаются на английском.

Visualising Relationships Between Variables

Distribution plots show one variable at a time. When you want to understand the relationship between two numeric variables, you need scatter plots and related visualisations. These help you detect correlation (do both variables increase together?), clusters (are there subgroups in the data?), and outliers (are there unusual points far from the main cloud?). Seaborn makes all of these easy with scatterplot and pairplot.

import seaborn as sns
import matplotlib.pyplot as plt

# Load the classic iris dataset
iris = sns.load_dataset('iris')
print(iris.head())
print('\nShape:', iris.shape)

Basic Scatter Plot with sns.scatterplot

sns.scatterplot(data, x, y) plots each observation as a point at its (x, y) coordinates. Unlike Matplotlib's plt.scatter, Seaborn's version automatically links to a DataFrame and integrates with hue, size, and style semantics. The result is informative multi-variable plots with a single function call and an automatic legend.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

sns.scatterplot(data=tips, x='total_bill', y='tip')
plt.title('Tip vs Total Bill')
plt.xlabel('Total Bill ($)')
plt.ylabel('Tip ($)')
plt.show()

Encoding a Third Variable with hue

The hue parameter assigns a colour to each point based on a third variable — either categorical or numeric. When hue is a categorical column (like 'smoker'), Seaborn uses distinct colours. When hue is a numeric column, it uses a sequential colour map. This lets you see three variables at once on a 2D plot without adding a third axis.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

sns.scatterplot(
    data=tips,
    x='total_bill',
    y='tip',
    hue='time',       # Lunch vs Dinner
    style='smoker',   # marker shape
    palette='deep'
)
plt.title('Tip vs Bill by Time and Smoker Status')
plt.legend(bbox_to_anchor=(1, 1))
plt.tight_layout()
plt.show()

Encoding Size and Style

Beyond hue, Seaborn scatter plots support size (point area encodes a numeric variable) and style (marker shape encodes a categorical variable). Using all three semantics simultaneously can reveal complex multi-variable patterns but risks overloading the viewer. A practical guideline is to use hue first (most salient), style second, and size only when necessary.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

sns.scatterplot(
    data=tips,
    x='total_bill',
    y='tip',
    hue='day',
    size='size',      # party size controls marker area
    sizes=(20, 200),  # min and max dot area
    alpha=0.7
)
plt.title('Tip vs Bill — Colour=Day, Size=Party Size')
plt.legend(bbox_to_anchor=(1, 1))
plt.tight_layout()
plt.show()

Adding a Regression Line with regplot

sns.regplot() draws a scatter plot and overlays a linear regression line with a shaded 95% confidence band. This is useful for visualising the strength and direction of the linear relationship between two variables. The confidence band narrows in regions with dense data and widens at the edges where fewer points inform the fit.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

sns.regplot(
    data=tips,
    x='total_bill',
    y='tip',
    scatter_kws={'alpha': 0.4},
    line_kws={'color': 'red'}
)
plt.title('Tip vs Bill with Regression Line')
plt.show()

lmplot: Faceted Regression Plots

sns.lmplot() is the figure-level version of regplot. It supports hue (separate regression lines per group in the same panel) and col/row (separate panels). This is powerful for testing whether the relationship between two variables differs across subgroups — for example, whether the bill-to-tip relationship is steeper for dinner than for lunch.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

# Separate regression line per time (Lunch / Dinner)
sns.lmplot(
    data=tips,
    x='total_bill',
    y='tip',
    hue='time',
    scatter_kws={'alpha': 0.4},
    height=5
)
plt.suptitle('Regression by Meal Time', y=1.02)
plt.show()

Introduction to Pair Plots

sns.pairplot(df) creates a grid of scatter plots for every pair of numeric columns in the DataFrame, with distribution plots on the diagonal. This is the fastest way to visualise all pairwise relationships in a dataset in one call. For a DataFrame with n numeric columns, pairplot creates an n×n grid. It is most practical when n is between 3 and 8 — larger grids become too small to read.

import seaborn as sns
import matplotlib.pyplot as plt

iris = sns.load_dataset('iris')

# Create pairplot coloured by species
sns.pairplot(iris, hue='species', diag_kind='kde', plot_kws={'alpha': 0.5})
plt.suptitle('Iris Dataset Pairplot', y=1.02)
plt.show()

Customising the Diagonal of pairplot

The diag_kind parameter controls what appears on the diagonal of the pairplot grid. Use 'hist' for histograms or 'kde' for kernel density estimates. The diagonal shows the univariate distribution of each variable, while the off-diagonal panels show pairwise scatter plots. You can also use corner=True to show only the lower triangle, halving the number of panels and reducing redundancy.

import seaborn as sns
import matplotlib.pyplot as plt

iris = sns.load_dataset('iris')

# Show only lower triangle
sns.pairplot(
    iris,
    hue='species',
    diag_kind='hist',
    corner=True
)
plt.show()

PairGrid for Full Customisation

sns.PairGrid gives you full control over which plot type appears in each section of the matrix. Use g.map_upper(), g.map_lower(), and g.map_diag() to assign different functions (like regplot on the upper triangle and kdeplot on the lower). This produces publication-quality plots where each panel conveys distinct information about the variable pair.

import seaborn as sns
import matplotlib.pyplot as plt

iris = sns.load_dataset('iris')
numeric_iris = iris.drop(columns=['species'])

g = sns.PairGrid(numeric_iris)
g.map_upper(sns.scatterplot, alpha=0.3)
g.map_lower(sns.kdeplot, fill=True)
g.map_diag(sns.histplot, kde=True)

plt.suptitle('Custom PairGrid', y=1.02)
plt.show()

Identifying Correlation in Scatter Plots

When reading a scatter plot, estimate the direction (does the cloud slope up or down?), strength (how tightly clustered are the points around the trend?), and form (linear or curved?) of the relationship. Correlation coefficient r ranges from -1 (perfect negative) to +1 (perfect positive), with 0 meaning no linear relationship. Remember: correlation is not causation, and a non-linear relationship can have r ≈ 0.

import seaborn as sns
import matplotlib.pyplot as plt
import numpy as np

np.random.seed(0)
x = np.linspace(0, 10, 100)

fig, axes = plt.subplots(1, 3, figsize=(12, 4))

# Strong positive correlation
axes[0].scatter(x, x + np.random.normal(0, 0.5, 100))
axes[0].set_title('Strong Positive (r≈0.99)')

# Weak correlation
axes[1].scatter(x, np.random.normal(5, 3, 100))
axes[1].set_title('No Correlation (r≈0)')

# Non-linear (quadratic) — r can be low
axes[2].scatter(x, (x-5)**2 + np.random.normal(0, 1, 100))
axes[2].set_title('Non-linear (r low despite pattern)')

plt.tight_layout()
plt.show()

jointplot for Bivariate with Marginals

sns.jointplot() combines a central bivariate plot with marginal univariate plots on the top and right edges. Pass kind='scatter', 'hex', 'kde', or 'reg'. The hex kind is useful when you have thousands of overlapping points — it bins them into hexagons and uses colour to show density, solving the overplotting problem that makes scatter plots unreadable for large datasets.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

# KDE joint plot with marginal distributions
sns.jointplot(
    data=tips,
    x='total_bill',
    y='tip',
    kind='reg',
    marginal_kws={'bins': 20}
)
plt.suptitle('Joint Distribution of Bill and Tip', y=1.02)
plt.show()

Quick Check

Test your understanding of Seaborn scatter and pair plots from this lesson.

Lesson Recap

In this lesson you learned: sns.scatterplot encodes up to four variables using x, y, hue, size, and style, sns.regplot/lmplot overlays a regression line to quantify linear relationships, and sns.pairplot produces an all-pairs grid ideal for exploratory analysis of multi-variable datasets. Next up we explore heatmaps for visualising correlation matrices.

Часто задаваемые вопросы

Урок «Диаграммы рассеяния и парные графики» бесплатный?

Да — полный текст урока «Диаграммы рассеяния и парные графики» бесплатно доступен здесь в веб-версии. Чтобы практиковать его интерактивно (встроенный редактор кода и ИИ-репетитор 24/7) и разблокировать остальной курс Pandas & NumPy Academy, подпишись на CoddyKit PRO. Курс Pandas & NumPy Academy содержит 4 уроков всего.

Чему я научусь в уроке «Диаграммы рассеяния и парные графики»?

Создавайте диаграммы рассеяния, окрашенные по третьей переменной, с помощью sns.scatterplot и визуализируйте все попарные взаимосвязи с помощью sns.pairplot. Ты практикуешь Pandas & NumPy Academy с помощью реального кода, который запускаешь прямо в браузере, и ИИ-репетитор 24/7 отвечает на твои вопросы во время урока.

Нужен ли мне опыт, чтобы начать Pandas & NumPy Academy?

Предыдущий опыт не требуется. Pandas & NumPy Academy на CoddyKit структурирован для всех уровней — от новичков до продвинутых, поэтому ты можешь начать отсюда или с самого начала и учиться в своем темпе. Это урок 3 из 4.

Сколько времени занимает урок «Диаграммы рассеяния и парные графики»?

Большинство уроков CoddyKit занимают около 5–10 минут. Каждый из них компактный и интерактивный, поэтому ты постоянно делаешь прогресс и продолжаешь с того же места в веб-версии и приложении.

Можно ли писать и запускать код в этом уроке Pandas & NumPy Academy?

Да. Каждый урок Pandas & NumPy Academy включает встроенный редактор кода, поэтому ты пишешь и запускаешь реальный код прямо в браузере и получаешь моментальную обратную связь от AI — локальная установка не требуется.

Все уроки этого курса

  1. Графики распределений: histplot и kdeplot
  2. Категориальные графики: boxplot, barplot, violinplot
  3. Диаграммы рассеяния и парные графики
  4. Тепловые карты корреляционных матриц
← Назад к Pandas & NumPy Academy