Machine Learning Academy · Leçon

Visualiser les données avec Matplotlib et Seaborn

Tracez des histogrammes, des nuages de points et des cartes de chaleur des corrélations pour explorer les distributions et les relations entre les données avant la modélisation.

Leçon 4 sur 413 étapes

Visualiser les données avec Matplotlib et Seaborn est une leçon Machine Learning Academy gratuite sur CoddyKit. Ceci est la leçon 4 sur 4. Tu peux lire la leçon complète ci-dessous gratuitement — puis la pratiquer en direct dans le navigateur avec un éditeur de code intégré et un tuteur IA 24/7. Elle fait partie du parcours d'apprentissage Machine Learning Academy, et ta progression se synchronise sur le web et l'application CoddyKit. Le cours Machine Learning Academy comprend 4 leçons au total.

Certaines parties de cette leçon n'ont pas encore été traduites et s'affichent en anglais.

Why Visualisation Matters in ML

Charts aren't just for reports — they're a diagnostic tool at every step. Matplotlib gives full control; Seaborn makes beautiful stats plots with less code.

import matplotlib.pyplot as plt
import seaborn as sns
import pandas as pd
import numpy as np

# Set seaborn theme for all plots
sns.set_theme(style='whitegrid', palette='muted')

# Load a built-in dataset
df = sns.load_dataset('tips')
print(df.head())

Histograms: Understanding Distributions

A histogram bins a numeric column to show its shape — normal, skewed, or with outliers. Spotting skew tells you when a log transform might help your model.

import matplotlib.pyplot as plt
import seaborn as sns

df = sns.load_dataset('tips')

fig, axes = plt.subplots(1, 2, figsize=(12, 4))

# Raw distribution (right-skewed)
axes[0].hist(df['total_bill'], bins=30, edgecolor='white')
axes[0].set_title('Total Bill Distribution (Raw)')

# After log transform
import numpy as np
axes[1].hist(np.log(df['total_bill']), bins=30, edgecolor='white')
axes[1].set_title('Total Bill Distribution (Log Transformed)')

plt.tight_layout()
plt.show()

Scatter Plots: Feature Relationships

A scatter plot shows how two numbers relate — linear, curved, or full of outliers. Add color with hue to squeeze a third variable into the same view.

import matplotlib.pyplot as plt
import seaborn as sns

df = sns.load_dataset('tips')

# Scatter plot with hue encoding
sns.scatterplot(
    data=df,
    x='total_bill',
    y='tip',
    hue='time',      # encode meal time as colour
    size='size',     # encode party size as dot size
    alpha=0.7
)
plt.title('Tip vs Total Bill (coloured by Meal Time)')
plt.show()

Box Plots: Comparing Groups

A box plot shows the median, spread, and outliers across groups. If a feature's box looks very different per category, that feature is probably worth keeping.

import matplotlib.pyplot as plt
import seaborn as sns

df = sns.load_dataset('tips')

fig, axes = plt.subplots(1, 2, figsize=(12, 5))

sns.boxplot(data=df, x='day', y='total_bill', ax=axes[0])
axes[0].set_title('Bill by Day of Week')

sns.boxplot(data=df, x='smoker', y='tip', hue='sex', ax=axes[1])
axes[1].set_title('Tip by Smoking Status and Sex')

plt.tight_layout()
plt.show()

Correlation Heatmap: Finding Related Features

A correlation heatmap colors how strongly every pair of columns moves together. It reveals good predictors of your target — and redundant features to drop.

import matplotlib.pyplot as plt
import seaborn as sns
import pandas as pd

df = pd.read_csv('titanic.csv')

# Compute correlation matrix
corr = df[['Survived', 'Pclass', 'Age', 'SibSp', 'Parch', 'Fare']].corr()

# Plot heatmap
plt.figure(figsize=(8, 6))
sns.heatmap(
    corr,
    annot=True,     # show correlation values
    fmt='.2f',      # 2 decimal places
    cmap='RdYlGn',  # red-yellow-green colour scale
    vmin=-1, vmax=1
)
plt.title('Feature Correlation Matrix')
plt.show()

Pair Plots: Exploring All Feature Pairs

A pair plot shows scatter plots for every feature pair at once — a fast first look at a new dataset. Color by class to see which features separate the groups.

import seaborn as sns
import matplotlib.pyplot as plt

# Use the iris dataset (classic ML benchmark)
df = sns.load_dataset('iris')

# Pair plot coloured by species
sns.pairplot(
    df,
    hue='species',
    diag_kind='kde',    # KDE on diagonal
    plot_kws={'alpha': 0.6}
)
plt.suptitle('Iris Dataset Pair Plot', y=1.02)
plt.show()

Bar Charts and Count Plots

A count plot shows how often each category appears — the quickest way to check class balance before training a classifier. Add hue to compare two categories.

import seaborn as sns
import matplotlib.pyplot as plt

df = sns.load_dataset('titanic')

fig, axes = plt.subplots(1, 2, figsize=(12, 5))

# Class distribution (check balance)
sns.countplot(data=df, x='survived', ax=axes[0])
axes[0].set_title('Survival Count')

# Class by passenger class and sex
sns.countplot(data=df, x='class', hue='sex', ax=axes[1])
axes[1].set_title('Class Distribution by Sex')

plt.tight_layout()
plt.show()

Plotting Learning Curves

A learning curve plots training vs validation score as data grows. Both low means underfitting; a big gap means overfitting; both high and close means a good fit.

import matplotlib.pyplot as plt
import numpy as np
from sklearn.model_selection import learning_curve
from sklearn.tree import DecisionTreeClassifier
from sklearn.datasets import load_breast_cancer

X, y = load_breast_cancer(return_X_y=True)
train_sizes, train_scores, val_scores = learning_curve(
    DecisionTreeClassifier(max_depth=5), X, y, cv=5
)

plt.plot(train_sizes, train_scores.mean(axis=1), label='Training Score')
plt.plot(train_sizes, val_scores.mean(axis=1), label='Validation Score')
plt.xlabel('Training Set Size')
plt.ylabel('Accuracy')
plt.legend()
plt.title('Learning Curve')
plt.show()

Visualising Model Predictions

After training, plot the predictions. A confusion matrix heatmap shows where a classifier confuses labels — far more telling than a single accuracy number.

import seaborn as sns
import matplotlib.pyplot as plt
from sklearn.metrics import confusion_matrix
from sklearn.datasets import load_iris
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split

X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)

model = DecisionTreeClassifier(max_depth=3)
model.fit(X_train, y_train)
y_pred = model.predict(X_test)

cm = confusion_matrix(y_test, y_pred)
sns.heatmap(cm, annot=True, fmt='d', cmap='Blues')
plt.xlabel('Predicted')
plt.ylabel('Actual')
plt.title('Confusion Matrix')
plt.show()

Saving and Customising Plots

Good plots need polish: titles, axis labels, and readable fonts. Save them with savefig at dpi=150+ for reports, and pick a colorblind-friendly palette.

import matplotlib.pyplot as plt
import seaborn as sns
import numpy as np

# Create figure and axes explicitly
fig, ax = plt.subplots(figsize=(8, 5))

x = np.linspace(0, 10, 100)
ax.plot(x, np.sin(x), label='sin(x)', linewidth=2)
ax.plot(x, np.cos(x), label='cos(x)', linewidth=2, linestyle='--')

# Customise
ax.set_xlabel('x', fontsize=13)
ax.set_ylabel('y', fontsize=13)
ax.set_title('Sine and Cosine', fontsize=15, fontweight='bold')
ax.legend(fontsize=12)
ax.grid(True, alpha=0.3)

# Save
fig.savefig('plot.png', dpi=150, bbox_inches='tight')
plt.show()

Distribution Plots with Seaborn

A violin plot blends a box plot with a density curve, showing the full shape of a distribution. Seaborn's displot and kdeplot are great for smooth comparisons.

import seaborn as sns
import matplotlib.pyplot as plt

df = sns.load_dataset('tips')

fig, axes = plt.subplots(1, 2, figsize=(12, 5))

# Violin plot
sns.violinplot(data=df, x='day', y='total_bill', hue='sex',
               split=True, inner='quart', ax=axes[0])
axes[0].set_title('Bill Distribution by Day (Violin)')

# KDE distribution comparison
sns.kdeplot(data=df, x='tip', hue='time', fill=True, alpha=0.4, ax=axes[1])
axes[1].set_title('Tip Distribution by Meal Time (KDE)')

plt.tight_layout()
plt.show()

Quick Check

Test your understanding of Machine Learning with Python concepts from this lesson.

Lesson Recap

You learned to see your data: histograms and scatter plots reveal shape, heatmaps find predictors, and learning curves diagnose fit. Next: your first model! 🚀

Gratuit pour commencer

Apprends Python avec un tuteur IA — gratuit

Écris et exécute du vrai code dans ton navigateur, obtiens de l'aide instantanée d'un tuteur IA disponible 24h/24, et reprends là où tu t'es arrêté sur le web ou dans l'app.

Cours
30
Leçons
120

Questions Fréquemment Posées

La leçon « Visualiser les données avec Matplotlib et Seaborn » est-elle gratuite ?

Oui — le texte complet de « Visualiser les données avec Matplotlib et Seaborn » est gratuit à lire ici sur le web. Pour la pratiquer de manière interactive (un éditeur de code intégré et un tuteur IA 24/7) et déverrouiller le reste du cours Machine Learning Academy, passe à CoddyKit PRO. Le cours Machine Learning Academy comprend 4 leçons au total.

Qu'est-ce que j'apprendrai dans « Visualiser les données avec Matplotlib et Seaborn » ?

Tracez des histogrammes, des nuages de points et des cartes de chaleur des corrélations pour explorer les distributions et les relations entre les données avant la modélisation. Tu pratiques Machine Learning Academy avec du code pratique que tu exécutes directement dans le navigateur, et un tuteur IA 24/7 répond à tes questions au fur et à mesure que tu avances dans la leçon.

Dois-je avoir de l'expérience pour commencer Machine Learning Academy ?

Aucune expérience préalable n'est requise. Machine Learning Academy sur CoddyKit est structuré pour les débutants jusqu'aux apprenants avancés, donc tu peux commencer ici ou depuis le début et avancer à ton rythme. Ceci est la leçon 4 sur 4.

Combien de temps prend la leçon « Visualiser les données avec Matplotlib et Seaborn » ?

La plupart des leçons CoddyKit prennent environ 5–10 minutes. Chacune est courte et interactive, tu progresses régulièrement et tu repiques exactement où tu t'es arrêté sur le web et l'app.

Peux-tu écrire et exécuter du code dans cette leçon Machine Learning Academy ?

Oui. Chaque leçon Machine Learning Academy inclut un éditeur de code intégré, tu écris et exécutes du vrai code directement dans ton navigateur et tu reçois des retours IA instantanés — aucune configuration locale requise.

Toutes les leçons de ce cours

  1. Installation d’Anaconda et de Jupyter Notebook
  2. Les fondamentaux de NumPy : tableaux et opérations mathématiques
  3. Pandas pour manipuler les données
  4. Visualiser les données avec Matplotlib et Seaborn
← Retour à Machine Learning Academy