0Pricing
Pandas & NumPy Academy · 课时

分布图:histplot 与 kdeplot

使用 sns.histplot 可视化数值分布的形状,并使用 sns.kdeplot 叠加核密度估计。

分布图:histplot 与 kdeplot 是 CoddyKit 上的免费 Pandas & NumPy Academy 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Pandas & NumPy Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Pandas & NumPy Academy 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Seaborn and Distribution Plots

Seaborn is a statistical data visualisation library built on top of Matplotlib. It provides a high-level API that makes beautiful, informative plots with minimal code. One of the first things you want to understand about any dataset is the shape of its distributions — and Seaborn's histplot and kdeplot are the primary tools for this.

import seaborn as sns
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd

# Load a built-in dataset
tips = sns.load_dataset('tips')
print(tips.head())

Creating a Basic Histogram with histplot

sns.histplot(data, x='column') creates a histogram of a numeric variable. Seaborn automatically chooses sensible bin counts using Sturges' rule by default. You can customise the number of bins with the bins parameter, or let Seaborn pick automatically. The height of each bar represents the count of observations in that range.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

# Basic histogram of total bill amounts
sns.histplot(data=tips, x='total_bill')
plt.title('Distribution of Total Bill')
plt.xlabel('Total Bill ($)')
plt.show()

Controlling Bins and Stat Parameter

The bins parameter controls how many bars the histogram shows — more bins reveal finer detail but can look noisy. The stat parameter changes what the y-axis represents: 'count' (default), 'density' (probability density), 'probability' (fraction of observations), or 'percent'. Choosing stat='density' makes histograms from different-sized datasets comparable.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

fig, axes = plt.subplots(1, 2, figsize=(10, 4))

# Count histogram with 20 bins
sns.histplot(data=tips, x='total_bill', bins=20, ax=axes[0])
axes[0].set_title('Count (20 bins)')

# Density histogram
sns.histplot(data=tips, x='total_bill', stat='density', ax=axes[1])
axes[1].set_title('Density')

plt.tight_layout()
plt.show()

Overlaying a KDE on the Histogram

Setting kde=True in histplot overlays a Kernel Density Estimate (KDE) curve on the histogram. The KDE is a smooth continuous approximation of the underlying probability distribution. This combination is very effective: the histogram shows the raw bin counts while the KDE reveals the overall shape and detects multiple peaks (multimodality).

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

# Histogram with KDE overlay
sns.histplot(data=tips, x='total_bill', kde=True, bins=25, color='steelblue')
plt.title('Bill Distribution with KDE Overlay')
plt.xlabel('Total Bill ($)')
plt.show()

Using kdeplot Independently

sns.kdeplot() draws only the smooth density curve without histogram bars. This is useful when comparing multiple distributions on the same axes because overlapping histograms become confusing, while overlapping KDE curves remain readable. The fill=True parameter shades the area under the curve, making it easier to distinguish groups visually.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

# KDE plot for lunch vs dinner bills
sns.kdeplot(data=tips, x='total_bill', hue='time', fill=True, alpha=0.5)
plt.title('Bill Distribution: Lunch vs Dinner')
plt.xlabel('Total Bill ($)')
plt.show()

The hue Parameter for Group Comparison

Both histplot and kdeplot accept a hue parameter that splits the data by a categorical column and draws each group in a different colour. This makes it easy to compare distributions across groups. When using histplot with hue, set stat='density' to make groups of different sizes fairly comparable.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

# Compare tip distributions by smoker status
sns.histplot(
    data=tips,
    x='tip',
    hue='smoker',
    stat='density',
    common_norm=False,
    bins=20,
    alpha=0.6
)
plt.title('Tip Distribution by Smoker Status')
plt.show()

Bandwidth in KDE: The bw_adjust Parameter

KDE has a tuning parameter called bandwidth that controls how smooth the curve is. A small bandwidth makes the curve jagged and overfits to the data; a large bandwidth over-smooths and hides real features. Seaborn uses bw_adjust (default 1.0) as a multiplier on the automatically chosen bandwidth — values below 1 increase roughness, above 1 increase smoothness.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

fig, axes = plt.subplots(1, 3, figsize=(12, 4))
bw_values = [0.3, 1.0, 3.0]

for ax, bw in zip(axes, bw_values):
    sns.kdeplot(data=tips, x='total_bill', bw_adjust=bw, ax=ax)
    ax.set_title(f'bw_adjust={bw}')

plt.tight_layout()
plt.show()

2D Distributions with histplot and kdeplot

Both histplot and kdeplot support two-dimensional plots by passing both x and y parameters. A 2D histogram shows a colour-coded frequency grid; a 2D KDE shows contour lines of equal density. These plots reveal the joint distribution of two numeric variables and whether they are correlated.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

fig, axes = plt.subplots(1, 2, figsize=(12, 5))

# 2D histogram
sns.histplot(data=tips, x='total_bill', y='tip', ax=axes[0])
axes[0].set_title('2D Histogram')

# 2D KDE contour plot
sns.kdeplot(data=tips, x='total_bill', y='tip', fill=True, ax=axes[1])
axes[1].set_title('2D KDE Contours')

plt.tight_layout()
plt.show()

Customising Appearance with Seaborn Themes

Seaborn provides built-in themes via sns.set_theme(style=). The available styles are 'darkgrid', 'whitegrid', 'dark', 'white', and 'ticks'. You can also set a palette with palette= or sns.set_palette(). These settings apply globally to all subsequent plots in the session, making it easy to achieve a consistent look.

import seaborn as sns
import matplotlib.pyplot as plt

# Apply a clean whitegrid theme
sns.set_theme(style='whitegrid', palette='muted')

tips = sns.load_dataset('tips')
sns.histplot(data=tips, x='total_bill', kde=True, bins=20)
plt.title('Styled Distribution Plot')
plt.show()

# Reset to defaults when done
sns.reset_defaults()

Interpreting Distribution Shape

When reading a distribution plot, look for four key features. Centre: where does the mass concentrate (mean/median)? Spread: how wide is the distribution (std dev)? Skewness: is the tail longer on the right (positive skew) or left (negative skew)? Modality: does the distribution have one peak (unimodal) or more (bimodal/multimodal)? Bimodal distributions often indicate two hidden subpopulations worth separating.

import seaborn as sns
import matplotlib.pyplot as plt
import numpy as np

# Simulate a bimodal distribution
np.random.seed(42)
data = np.concatenate([
    np.random.normal(loc=20, scale=3, size=300),
    np.random.normal(loc=45, scale=5, size=200)
])

sns.histplot(data, kde=True, bins=40)
plt.title('Bimodal Distribution — Two Hidden Groups')
plt.xlabel('Value')
plt.show()

displot: Figure-Level Distribution Function

sns.displot() is the figure-level wrapper that creates its own Figure and supports facetting across rows and columns using col= and row= parameters. Pass kind='hist', kind='kde', or kind='ecdf' to switch plot types. The ECDF (Empirical Cumulative Distribution Function) is especially useful because it shows what fraction of the data lies below each value without requiring bin choices.

import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset('tips')

# Faceted KDE by day of week
sns.displot(
    data=tips,
    x='total_bill',
    col='day',
    col_wrap=2,
    kind='kde',
    fill=True
)
plt.suptitle('Bill Distributions by Day', y=1.02)
plt.show()

Quick Check

Test your understanding of Seaborn distribution plots from this lesson.

Lesson Recap

In this lesson you learned: sns.histplot creates histograms with customisable bins and stat parameters, sns.kdeplot draws smooth density curves ideal for group comparisons, and the hue parameter splits both plot types by a categorical variable for side-by-side comparison. Next up we explore categorical comparison plots like box plots, bar plots, and violin plots.

常见问题解答

「分布图:histplot 与 kdeplot」课时是免费的吗?

是的 — 「分布图:histplot 与 kdeplot」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Pandas & NumPy Academy 课程的其余内容,请升级到 CoddyKit PRO。 Pandas & NumPy Academy 课程共包含 4 节课。

「分布图:histplot 与 kdeplot」这节课中我会学到什么?

使用 sns.histplot 可视化数值分布的形状,并使用 sns.kdeplot 叠加核密度估计。 你通过在浏览器中直接运行的动手代码来练习 Pandas & NumPy Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Pandas & NumPy Academy 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Pandas & NumPy Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。

「分布图:histplot 与 kdeplot」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Pandas & NumPy Academy 课中编写并运行代码吗?

能。每节 Pandas & NumPy Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 分布图:histplot 与 kdeplot
  2. 分类图:boxplot、barplot、violinplot
  3. 散点图与成对关系图
  4. 相关矩阵热力图
← 返回 Pandas & NumPy Academy