Graphiques de distribution : histplot et kdeplot
Visualisez la forme d’une distribution numérique avec sns.histplot et superposez une estimation de densité par noyau avec sns.kdeplot.
Graphiques de distribution : histplot et kdeplot est une leçon Pandas & NumPy Academy gratuite sur CoddyKit. Ceci est la leçon 1 sur 4. Tu peux lire la leçon complète ci-dessous gratuitement — puis la pratiquer en direct dans le navigateur avec un éditeur de code intégré et un tuteur IA 24/7. Elle fait partie du parcours d'apprentissage Pandas & NumPy Academy, et ta progression se synchronise sur le web et l'application CoddyKit. Le cours Pandas & NumPy Academy comprend 4 leçons au total.
Certaines parties de cette leçon n'ont pas encore été traduites et s'affichent en anglais.
Seaborn and Distribution Plots
Seaborn is a statistical data visualisation library built on top of Matplotlib. It provides a high-level API that makes beautiful, informative plots with minimal code. One of the first things you want to understand about any dataset is the shape of its distributions — and Seaborn's histplot and kdeplot are the primary tools for this.
import seaborn as sns
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
# Load a built-in dataset
tips = sns.load_dataset('tips')
print(tips.head())Creating a Basic Histogram with histplot
sns.histplot(data, x='column') creates a histogram of a numeric variable. Seaborn automatically chooses sensible bin counts using Sturges' rule by default. You can customise the number of bins with the bins parameter, or let Seaborn pick automatically. The height of each bar represents the count of observations in that range.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
# Basic histogram of total bill amounts
sns.histplot(data=tips, x='total_bill')
plt.title('Distribution of Total Bill')
plt.xlabel('Total Bill ($)')
plt.show()Controlling Bins and Stat Parameter
The bins parameter controls how many bars the histogram shows — more bins reveal finer detail but can look noisy. The stat parameter changes what the y-axis represents: 'count' (default), 'density' (probability density), 'probability' (fraction of observations), or 'percent'. Choosing stat='density' makes histograms from different-sized datasets comparable.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
fig, axes = plt.subplots(1, 2, figsize=(10, 4))
# Count histogram with 20 bins
sns.histplot(data=tips, x='total_bill', bins=20, ax=axes[0])
axes[0].set_title('Count (20 bins)')
# Density histogram
sns.histplot(data=tips, x='total_bill', stat='density', ax=axes[1])
axes[1].set_title('Density')
plt.tight_layout()
plt.show()Overlaying a KDE on the Histogram
Setting kde=True in histplot overlays a Kernel Density Estimate (KDE) curve on the histogram. The KDE is a smooth continuous approximation of the underlying probability distribution. This combination is very effective: the histogram shows the raw bin counts while the KDE reveals the overall shape and detects multiple peaks (multimodality).
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
# Histogram with KDE overlay
sns.histplot(data=tips, x='total_bill', kde=True, bins=25, color='steelblue')
plt.title('Bill Distribution with KDE Overlay')
plt.xlabel('Total Bill ($)')
plt.show()Using kdeplot Independently
sns.kdeplot() draws only the smooth density curve without histogram bars. This is useful when comparing multiple distributions on the same axes because overlapping histograms become confusing, while overlapping KDE curves remain readable. The fill=True parameter shades the area under the curve, making it easier to distinguish groups visually.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
# KDE plot for lunch vs dinner bills
sns.kdeplot(data=tips, x='total_bill', hue='time', fill=True, alpha=0.5)
plt.title('Bill Distribution: Lunch vs Dinner')
plt.xlabel('Total Bill ($)')
plt.show()The hue Parameter for Group Comparison
Both histplot and kdeplot accept a hue parameter that splits the data by a categorical column and draws each group in a different colour. This makes it easy to compare distributions across groups. When using histplot with hue, set stat='density' to make groups of different sizes fairly comparable.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
# Compare tip distributions by smoker status
sns.histplot(
data=tips,
x='tip',
hue='smoker',
stat='density',
common_norm=False,
bins=20,
alpha=0.6
)
plt.title('Tip Distribution by Smoker Status')
plt.show()Bandwidth in KDE: The bw_adjust Parameter
KDE has a tuning parameter called bandwidth that controls how smooth the curve is. A small bandwidth makes the curve jagged and overfits to the data; a large bandwidth over-smooths and hides real features. Seaborn uses bw_adjust (default 1.0) as a multiplier on the automatically chosen bandwidth — values below 1 increase roughness, above 1 increase smoothness.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
fig, axes = plt.subplots(1, 3, figsize=(12, 4))
bw_values = [0.3, 1.0, 3.0]
for ax, bw in zip(axes, bw_values):
sns.kdeplot(data=tips, x='total_bill', bw_adjust=bw, ax=ax)
ax.set_title(f'bw_adjust={bw}')
plt.tight_layout()
plt.show()2D Distributions with histplot and kdeplot
Both histplot and kdeplot support two-dimensional plots by passing both x and y parameters. A 2D histogram shows a colour-coded frequency grid; a 2D KDE shows contour lines of equal density. These plots reveal the joint distribution of two numeric variables and whether they are correlated.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
fig, axes = plt.subplots(1, 2, figsize=(12, 5))
# 2D histogram
sns.histplot(data=tips, x='total_bill', y='tip', ax=axes[0])
axes[0].set_title('2D Histogram')
# 2D KDE contour plot
sns.kdeplot(data=tips, x='total_bill', y='tip', fill=True, ax=axes[1])
axes[1].set_title('2D KDE Contours')
plt.tight_layout()
plt.show()Customising Appearance with Seaborn Themes
Seaborn provides built-in themes via sns.set_theme(style=). The available styles are 'darkgrid', 'whitegrid', 'dark', 'white', and 'ticks'. You can also set a palette with palette= or sns.set_palette(). These settings apply globally to all subsequent plots in the session, making it easy to achieve a consistent look.
import seaborn as sns
import matplotlib.pyplot as plt
# Apply a clean whitegrid theme
sns.set_theme(style='whitegrid', palette='muted')
tips = sns.load_dataset('tips')
sns.histplot(data=tips, x='total_bill', kde=True, bins=20)
plt.title('Styled Distribution Plot')
plt.show()
# Reset to defaults when done
sns.reset_defaults()Interpreting Distribution Shape
When reading a distribution plot, look for four key features. Centre: where does the mass concentrate (mean/median)? Spread: how wide is the distribution (std dev)? Skewness: is the tail longer on the right (positive skew) or left (negative skew)? Modality: does the distribution have one peak (unimodal) or more (bimodal/multimodal)? Bimodal distributions often indicate two hidden subpopulations worth separating.
import seaborn as sns
import matplotlib.pyplot as plt
import numpy as np
# Simulate a bimodal distribution
np.random.seed(42)
data = np.concatenate([
np.random.normal(loc=20, scale=3, size=300),
np.random.normal(loc=45, scale=5, size=200)
])
sns.histplot(data, kde=True, bins=40)
plt.title('Bimodal Distribution — Two Hidden Groups')
plt.xlabel('Value')
plt.show()displot: Figure-Level Distribution Function
sns.displot() is the figure-level wrapper that creates its own Figure and supports facetting across rows and columns using col= and row= parameters. Pass kind='hist', kind='kde', or kind='ecdf' to switch plot types. The ECDF (Empirical Cumulative Distribution Function) is especially useful because it shows what fraction of the data lies below each value without requiring bin choices.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
# Faceted KDE by day of week
sns.displot(
data=tips,
x='total_bill',
col='day',
col_wrap=2,
kind='kde',
fill=True
)
plt.suptitle('Bill Distributions by Day', y=1.02)
plt.show()Quick Check
Test your understanding of Seaborn distribution plots from this lesson.
Lesson Recap
In this lesson you learned: sns.histplot creates histograms with customisable bins and stat parameters, sns.kdeplot draws smooth density curves ideal for group comparisons, and the hue parameter splits both plot types by a categorical variable for side-by-side comparison. Next up we explore categorical comparison plots like box plots, bar plots, and violin plots.
Questions Fréquemment Posées
La leçon « Graphiques de distribution : histplot et kdeplot » est-elle gratuite ?
Oui — le texte complet de « Graphiques de distribution : histplot et kdeplot » est gratuit à lire ici sur le web. Pour la pratiquer de manière interactive (un éditeur de code intégré et un tuteur IA 24/7) et déverrouiller le reste du cours Pandas & NumPy Academy, passe à CoddyKit PRO. Le cours Pandas & NumPy Academy comprend 4 leçons au total.
Qu'est-ce que j'apprendrai dans « Graphiques de distribution : histplot et kdeplot » ?
Visualisez la forme d’une distribution numérique avec sns.histplot et superposez une estimation de densité par noyau avec sns.kdeplot. Tu pratiques Pandas & NumPy Academy avec du code pratique que tu exécutes directement dans le navigateur, et un tuteur IA 24/7 répond à tes questions au fur et à mesure que tu avances dans la leçon.
Dois-je avoir de l'expérience pour commencer Pandas & NumPy Academy ?
Aucune expérience préalable n'est requise. Pandas & NumPy Academy sur CoddyKit est structuré pour les débutants jusqu'aux apprenants avancés, donc tu peux commencer ici ou depuis le début et avancer à ton rythme. Ceci est la leçon 1 sur 4.
Combien de temps prend la leçon « Graphiques de distribution : histplot et kdeplot » ?
La plupart des leçons CoddyKit prennent environ 5–10 minutes. Chacune est courte et interactive, tu progresses régulièrement et tu repiques exactement où tu t'es arrêté sur le web et l'app.
Peux-tu écrire et exécuter du code dans cette leçon Pandas & NumPy Academy ?
Oui. Chaque leçon Pandas & NumPy Academy inclut un éditeur de code intégré, tu écris et exécutes du vrai code directement dans ton navigateur et tu reçois des retours IA instantanés — aucune configuration locale requise.
Toutes les leçons de ce cours
- Graphiques de distribution : histplot et kdeplot
- Graphiques catégoriels : boxplot, barplot, violinplot
- Nuages de points et graphiques de paires
- Cartes thermiques des matrices de corrélation