0Pricing
Pandas & NumPy Academy · Leçon

Nuages de points et graphiques de paires

Créez des nuages de points colorés selon une troisième variable avec sns.scatterplot et visualisez toutes les relations deux à deux avec sns.pairplot.

Nuages de points et graphiques de paires est une leçon Pandas & NumPy Academy gratuite sur CoddyKit. Ceci est la leçon 3 sur 4. Tu peux lire la leçon complète ci-dessous gratuitement — puis la pratiquer en direct dans le navigateur avec un éditeur de code intégré et un tuteur IA 24/7. Elle fait partie du parcours d'apprentissage Pandas & NumPy Academy, et ta progression se synchronise sur le web et l'application CoddyKit. Le cours Pandas & NumPy Academy comprend 4 leçons au total.

Certaines parties de cette leçon n'ont pas encore été traduites et s'affichent en anglais.

Visualising Relationships Between Variables

Distribution plots show one variable at a time. When you want to understand the relationship between two numeric variables, you need scatter plots and related visualisations. These help you detect correlation (do both variables increase together?), clusters (are there subgroups in the data?), and outliers (are there unusual points far from the main cloud?). Seaborn makes all of these easy with scatterplot and pairplot.

import seaborn as sns
import matplotlib.pyplot as plt

# Load the classic iris dataset
iris = sns.load_dataset('iris')
print(iris.head())
print('\nShape:', iris.shape)

Basic Scatter Plot with sns.scatterplot

sns.scatterplot(data, x, y) plots each observation as a point at its (x, y) coordinates. Unlike Matplotlib's plt.scatter, Seaborn's version automatically links to a DataFrame and integrates with hue, size, and style semantics. The result is informative multi-variable plots with a single function call and an automatic legend.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

sns.scatterplot(data=tips, x='total_bill', y='tip')
plt.title('Tip vs Total Bill')
plt.xlabel('Total Bill ($)')
plt.ylabel('Tip ($)')
plt.show()

Encoding a Third Variable with hue

The hue parameter assigns a colour to each point based on a third variable — either categorical or numeric. When hue is a categorical column (like 'smoker'), Seaborn uses distinct colours. When hue is a numeric column, it uses a sequential colour map. This lets you see three variables at once on a 2D plot without adding a third axis.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

sns.scatterplot(
    data=tips,
    x='total_bill',
    y='tip',
    hue='time',       # Lunch vs Dinner
    style='smoker',   # marker shape
    palette='deep'
)
plt.title('Tip vs Bill by Time and Smoker Status')
plt.legend(bbox_to_anchor=(1, 1))
plt.tight_layout()
plt.show()

Encoding Size and Style

Beyond hue, Seaborn scatter plots support size (point area encodes a numeric variable) and style (marker shape encodes a categorical variable). Using all three semantics simultaneously can reveal complex multi-variable patterns but risks overloading the viewer. A practical guideline is to use hue first (most salient), style second, and size only when necessary.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

sns.scatterplot(
    data=tips,
    x='total_bill',
    y='tip',
    hue='day',
    size='size',      # party size controls marker area
    sizes=(20, 200),  # min and max dot area
    alpha=0.7
)
plt.title('Tip vs Bill — Colour=Day, Size=Party Size')
plt.legend(bbox_to_anchor=(1, 1))
plt.tight_layout()
plt.show()

Adding a Regression Line with regplot

sns.regplot() draws a scatter plot and overlays a linear regression line with a shaded 95% confidence band. This is useful for visualising the strength and direction of the linear relationship between two variables. The confidence band narrows in regions with dense data and widens at the edges where fewer points inform the fit.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

sns.regplot(
    data=tips,
    x='total_bill',
    y='tip',
    scatter_kws={'alpha': 0.4},
    line_kws={'color': 'red'}
)
plt.title('Tip vs Bill with Regression Line')
plt.show()

lmplot: Faceted Regression Plots

sns.lmplot() is the figure-level version of regplot. It supports hue (separate regression lines per group in the same panel) and col/row (separate panels). This is powerful for testing whether the relationship between two variables differs across subgroups — for example, whether the bill-to-tip relationship is steeper for dinner than for lunch.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

# Separate regression line per time (Lunch / Dinner)
sns.lmplot(
    data=tips,
    x='total_bill',
    y='tip',
    hue='time',
    scatter_kws={'alpha': 0.4},
    height=5
)
plt.suptitle('Regression by Meal Time', y=1.02)
plt.show()

Introduction to Pair Plots

sns.pairplot(df) creates a grid of scatter plots for every pair of numeric columns in the DataFrame, with distribution plots on the diagonal. This is the fastest way to visualise all pairwise relationships in a dataset in one call. For a DataFrame with n numeric columns, pairplot creates an n×n grid. It is most practical when n is between 3 and 8 — larger grids become too small to read.

import seaborn as sns
import matplotlib.pyplot as plt

iris = sns.load_dataset('iris')

# Create pairplot coloured by species
sns.pairplot(iris, hue='species', diag_kind='kde', plot_kws={'alpha': 0.5})
plt.suptitle('Iris Dataset Pairplot', y=1.02)
plt.show()

Customising the Diagonal of pairplot

The diag_kind parameter controls what appears on the diagonal of the pairplot grid. Use 'hist' for histograms or 'kde' for kernel density estimates. The diagonal shows the univariate distribution of each variable, while the off-diagonal panels show pairwise scatter plots. You can also use corner=True to show only the lower triangle, halving the number of panels and reducing redundancy.

import seaborn as sns
import matplotlib.pyplot as plt

iris = sns.load_dataset('iris')

# Show only lower triangle
sns.pairplot(
    iris,
    hue='species',
    diag_kind='hist',
    corner=True
)
plt.show()

PairGrid for Full Customisation

sns.PairGrid gives you full control over which plot type appears in each section of the matrix. Use g.map_upper(), g.map_lower(), and g.map_diag() to assign different functions (like regplot on the upper triangle and kdeplot on the lower). This produces publication-quality plots where each panel conveys distinct information about the variable pair.

import seaborn as sns
import matplotlib.pyplot as plt

iris = sns.load_dataset('iris')
numeric_iris = iris.drop(columns=['species'])

g = sns.PairGrid(numeric_iris)
g.map_upper(sns.scatterplot, alpha=0.3)
g.map_lower(sns.kdeplot, fill=True)
g.map_diag(sns.histplot, kde=True)

plt.suptitle('Custom PairGrid', y=1.02)
plt.show()

Identifying Correlation in Scatter Plots

When reading a scatter plot, estimate the direction (does the cloud slope up or down?), strength (how tightly clustered are the points around the trend?), and form (linear or curved?) of the relationship. Correlation coefficient r ranges from -1 (perfect negative) to +1 (perfect positive), with 0 meaning no linear relationship. Remember: correlation is not causation, and a non-linear relationship can have r ≈ 0.

import seaborn as sns
import matplotlib.pyplot as plt
import numpy as np

np.random.seed(0)
x = np.linspace(0, 10, 100)

fig, axes = plt.subplots(1, 3, figsize=(12, 4))

# Strong positive correlation
axes[0].scatter(x, x + np.random.normal(0, 0.5, 100))
axes[0].set_title('Strong Positive (r≈0.99)')

# Weak correlation
axes[1].scatter(x, np.random.normal(5, 3, 100))
axes[1].set_title('No Correlation (r≈0)')

# Non-linear (quadratic) — r can be low
axes[2].scatter(x, (x-5)**2 + np.random.normal(0, 1, 100))
axes[2].set_title('Non-linear (r low despite pattern)')

plt.tight_layout()
plt.show()

jointplot for Bivariate with Marginals

sns.jointplot() combines a central bivariate plot with marginal univariate plots on the top and right edges. Pass kind='scatter', 'hex', 'kde', or 'reg'. The hex kind is useful when you have thousands of overlapping points — it bins them into hexagons and uses colour to show density, solving the overplotting problem that makes scatter plots unreadable for large datasets.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

# KDE joint plot with marginal distributions
sns.jointplot(
    data=tips,
    x='total_bill',
    y='tip',
    kind='reg',
    marginal_kws={'bins': 20}
)
plt.suptitle('Joint Distribution of Bill and Tip', y=1.02)
plt.show()

Quick Check

Test your understanding of Seaborn scatter and pair plots from this lesson.

Lesson Recap

In this lesson you learned: sns.scatterplot encodes up to four variables using x, y, hue, size, and style, sns.regplot/lmplot overlays a regression line to quantify linear relationships, and sns.pairplot produces an all-pairs grid ideal for exploratory analysis of multi-variable datasets. Next up we explore heatmaps for visualising correlation matrices.

Questions Fréquemment Posées

La leçon « Nuages de points et graphiques de paires » est-elle gratuite ?

Oui — le texte complet de « Nuages de points et graphiques de paires » est gratuit à lire ici sur le web. Pour la pratiquer de manière interactive (un éditeur de code intégré et un tuteur IA 24/7) et déverrouiller le reste du cours Pandas & NumPy Academy, passe à CoddyKit PRO. Le cours Pandas & NumPy Academy comprend 4 leçons au total.

Qu'est-ce que j'apprendrai dans « Nuages de points et graphiques de paires » ?

Créez des nuages de points colorés selon une troisième variable avec sns.scatterplot et visualisez toutes les relations deux à deux avec sns.pairplot. Tu pratiques Pandas & NumPy Academy avec du code pratique que tu exécutes directement dans le navigateur, et un tuteur IA 24/7 répond à tes questions au fur et à mesure que tu avances dans la leçon.

Dois-je avoir de l'expérience pour commencer Pandas & NumPy Academy ?

Aucune expérience préalable n'est requise. Pandas & NumPy Academy sur CoddyKit est structuré pour les débutants jusqu'aux apprenants avancés, donc tu peux commencer ici ou depuis le début et avancer à ton rythme. Ceci est la leçon 3 sur 4.

Combien de temps prend la leçon « Nuages de points et graphiques de paires » ?

La plupart des leçons CoddyKit prennent environ 5–10 minutes. Chacune est courte et interactive, tu progresses régulièrement et tu repiques exactement où tu t'es arrêté sur le web et l'app.

Peux-tu écrire et exécuter du code dans cette leçon Pandas & NumPy Academy ?

Oui. Chaque leçon Pandas & NumPy Academy inclut un éditeur de code intégré, tu écris et exécutes du vrai code directement dans ton navigateur et tu reçois des retours IA instantanés — aucune configuration locale requise.

Toutes les leçons de ce cours

  1. Graphiques de distribution : histplot et kdeplot
  2. Graphiques catégoriels : boxplot, barplot, violinplot
  3. Nuages de points et graphiques de paires
  4. Cartes thermiques des matrices de corrélation
← Retour à Pandas & NumPy Academy