Verteilungsdiagramme: histplot und kdeplot
Visualisieren Sie die Form einer numerischen Verteilung mit sns.histplot und legen Sie mit sns.kdeplot eine Kerndichteschätzung darüber.
Verteilungsdiagramme: histplot und kdeplot ist eine kostenlose Pandas & NumPy Academy-Lektion auf CoddyKit. Dies ist Lektion 1 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des Pandas & NumPy Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der Pandas & NumPy Academy-Kurs umfasst insgesamt 4 Lektionen.
Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.
Seaborn and Distribution Plots
Seaborn is a statistical data visualisation library built on top of Matplotlib. It provides a high-level API that makes beautiful, informative plots with minimal code. One of the first things you want to understand about any dataset is the shape of its distributions — and Seaborn's histplot and kdeplot are the primary tools for this.
import seaborn as sns
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
# Load a built-in dataset
tips = sns.load_dataset('tips')
print(tips.head())Creating a Basic Histogram with histplot
sns.histplot(data, x='column') creates a histogram of a numeric variable. Seaborn automatically chooses sensible bin counts using Sturges' rule by default. You can customise the number of bins with the bins parameter, or let Seaborn pick automatically. The height of each bar represents the count of observations in that range.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
# Basic histogram of total bill amounts
sns.histplot(data=tips, x='total_bill')
plt.title('Distribution of Total Bill')
plt.xlabel('Total Bill ($)')
plt.show()Controlling Bins and Stat Parameter
The bins parameter controls how many bars the histogram shows — more bins reveal finer detail but can look noisy. The stat parameter changes what the y-axis represents: 'count' (default), 'density' (probability density), 'probability' (fraction of observations), or 'percent'. Choosing stat='density' makes histograms from different-sized datasets comparable.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
fig, axes = plt.subplots(1, 2, figsize=(10, 4))
# Count histogram with 20 bins
sns.histplot(data=tips, x='total_bill', bins=20, ax=axes[0])
axes[0].set_title('Count (20 bins)')
# Density histogram
sns.histplot(data=tips, x='total_bill', stat='density', ax=axes[1])
axes[1].set_title('Density')
plt.tight_layout()
plt.show()Overlaying a KDE on the Histogram
Setting kde=True in histplot overlays a Kernel Density Estimate (KDE) curve on the histogram. The KDE is a smooth continuous approximation of the underlying probability distribution. This combination is very effective: the histogram shows the raw bin counts while the KDE reveals the overall shape and detects multiple peaks (multimodality).
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
# Histogram with KDE overlay
sns.histplot(data=tips, x='total_bill', kde=True, bins=25, color='steelblue')
plt.title('Bill Distribution with KDE Overlay')
plt.xlabel('Total Bill ($)')
plt.show()Using kdeplot Independently
sns.kdeplot() draws only the smooth density curve without histogram bars. This is useful when comparing multiple distributions on the same axes because overlapping histograms become confusing, while overlapping KDE curves remain readable. The fill=True parameter shades the area under the curve, making it easier to distinguish groups visually.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
# KDE plot for lunch vs dinner bills
sns.kdeplot(data=tips, x='total_bill', hue='time', fill=True, alpha=0.5)
plt.title('Bill Distribution: Lunch vs Dinner')
plt.xlabel('Total Bill ($)')
plt.show()The hue Parameter for Group Comparison
Both histplot and kdeplot accept a hue parameter that splits the data by a categorical column and draws each group in a different colour. This makes it easy to compare distributions across groups. When using histplot with hue, set stat='density' to make groups of different sizes fairly comparable.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
# Compare tip distributions by smoker status
sns.histplot(
data=tips,
x='tip',
hue='smoker',
stat='density',
common_norm=False,
bins=20,
alpha=0.6
)
plt.title('Tip Distribution by Smoker Status')
plt.show()Bandwidth in KDE: The bw_adjust Parameter
KDE has a tuning parameter called bandwidth that controls how smooth the curve is. A small bandwidth makes the curve jagged and overfits to the data; a large bandwidth over-smooths and hides real features. Seaborn uses bw_adjust (default 1.0) as a multiplier on the automatically chosen bandwidth — values below 1 increase roughness, above 1 increase smoothness.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
fig, axes = plt.subplots(1, 3, figsize=(12, 4))
bw_values = [0.3, 1.0, 3.0]
for ax, bw in zip(axes, bw_values):
sns.kdeplot(data=tips, x='total_bill', bw_adjust=bw, ax=ax)
ax.set_title(f'bw_adjust={bw}')
plt.tight_layout()
plt.show()2D Distributions with histplot and kdeplot
Both histplot and kdeplot support two-dimensional plots by passing both x and y parameters. A 2D histogram shows a colour-coded frequency grid; a 2D KDE shows contour lines of equal density. These plots reveal the joint distribution of two numeric variables and whether they are correlated.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
fig, axes = plt.subplots(1, 2, figsize=(12, 5))
# 2D histogram
sns.histplot(data=tips, x='total_bill', y='tip', ax=axes[0])
axes[0].set_title('2D Histogram')
# 2D KDE contour plot
sns.kdeplot(data=tips, x='total_bill', y='tip', fill=True, ax=axes[1])
axes[1].set_title('2D KDE Contours')
plt.tight_layout()
plt.show()Customising Appearance with Seaborn Themes
Seaborn provides built-in themes via sns.set_theme(style=). The available styles are 'darkgrid', 'whitegrid', 'dark', 'white', and 'ticks'. You can also set a palette with palette= or sns.set_palette(). These settings apply globally to all subsequent plots in the session, making it easy to achieve a consistent look.
import seaborn as sns
import matplotlib.pyplot as plt
# Apply a clean whitegrid theme
sns.set_theme(style='whitegrid', palette='muted')
tips = sns.load_dataset('tips')
sns.histplot(data=tips, x='total_bill', kde=True, bins=20)
plt.title('Styled Distribution Plot')
plt.show()
# Reset to defaults when done
sns.reset_defaults()Interpreting Distribution Shape
When reading a distribution plot, look for four key features. Centre: where does the mass concentrate (mean/median)? Spread: how wide is the distribution (std dev)? Skewness: is the tail longer on the right (positive skew) or left (negative skew)? Modality: does the distribution have one peak (unimodal) or more (bimodal/multimodal)? Bimodal distributions often indicate two hidden subpopulations worth separating.
import seaborn as sns
import matplotlib.pyplot as plt
import numpy as np
# Simulate a bimodal distribution
np.random.seed(42)
data = np.concatenate([
np.random.normal(loc=20, scale=3, size=300),
np.random.normal(loc=45, scale=5, size=200)
])
sns.histplot(data, kde=True, bins=40)
plt.title('Bimodal Distribution — Two Hidden Groups')
plt.xlabel('Value')
plt.show()displot: Figure-Level Distribution Function
sns.displot() is the figure-level wrapper that creates its own Figure and supports facetting across rows and columns using col= and row= parameters. Pass kind='hist', kind='kde', or kind='ecdf' to switch plot types. The ECDF (Empirical Cumulative Distribution Function) is especially useful because it shows what fraction of the data lies below each value without requiring bin choices.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
# Faceted KDE by day of week
sns.displot(
data=tips,
x='total_bill',
col='day',
col_wrap=2,
kind='kde',
fill=True
)
plt.suptitle('Bill Distributions by Day', y=1.02)
plt.show()Quick Check
Test your understanding of Seaborn distribution plots from this lesson.
Lesson Recap
In this lesson you learned: sns.histplot creates histograms with customisable bins and stat parameters, sns.kdeplot draws smooth density curves ideal for group comparisons, and the hue parameter splits both plot types by a categorical variable for side-by-side comparison. Next up we explore categorical comparison plots like box plots, bar plots, and violin plots.
Häufig gestellte Fragen
Ist die Lektion „Verteilungsdiagramme: histplot und kdeplot“ kostenlos?
Ja — der vollständige Text von „Verteilungsdiagramme: histplot und kdeplot“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des Pandas & NumPy Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der Pandas & NumPy Academy-Kurs umfasst insgesamt 4 Lektionen.
Was lerne ich in „Verteilungsdiagramme: histplot und kdeplot“?
Visualisieren Sie die Form einer numerischen Verteilung mit sns.histplot und legen Sie mit sns.kdeplot eine Kerndichteschätzung darüber. Du übst Pandas & NumPy Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.
Brauche ich Erfahrung, um Pandas & NumPy Academy zu starten?
Keine Vorkenntnisse erforderlich. Pandas & NumPy Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 1 von 4.
Wie lange dauert die Lektion „Verteilungsdiagramme: histplot und kdeplot“?
Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.
Kann ich in dieser Pandas & NumPy Academy-Lektion Code schreiben und ausführen?
Ja. Jede Pandas & NumPy Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.
Alle Lektionen in diesem Kurs
- Verteilungsdiagramme: histplot und kdeplot
- Kategoriale Diagramme: boxplot, barplot, violinplot
- Streudiagramme und Pairplots
- Heatmaps für Korrelationsmatrizen