散布図とペアプロット
sns.scatterplotで第3の変数に応じて色分けした散布図を作成し、sns.pairplotですべてのペア間の関係を可視化します。
「散布図とペアプロット」はCoddyKit上の無料Pandas & NumPy Academyレッスンです。 これはレッスン3/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはPandas & NumPy Academy学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Pandas & NumPy Academyコースには全4レッスンが含まれています。
このレッスンの一部はまだ翻訳されておらず、英語で表示されています。
Visualising Relationships Between Variables
Distribution plots show one variable at a time. When you want to understand the relationship between two numeric variables, you need scatter plots and related visualisations. These help you detect correlation (do both variables increase together?), clusters (are there subgroups in the data?), and outliers (are there unusual points far from the main cloud?). Seaborn makes all of these easy with scatterplot and pairplot.
import seaborn as sns
import matplotlib.pyplot as plt
# Load the classic iris dataset
iris = sns.load_dataset('iris')
print(iris.head())
print('\nShape:', iris.shape)Basic Scatter Plot with sns.scatterplot
sns.scatterplot(data, x, y) plots each observation as a point at its (x, y) coordinates. Unlike Matplotlib's plt.scatter, Seaborn's version automatically links to a DataFrame and integrates with hue, size, and style semantics. The result is informative multi-variable plots with a single function call and an automatic legend.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
sns.scatterplot(data=tips, x='total_bill', y='tip')
plt.title('Tip vs Total Bill')
plt.xlabel('Total Bill ($)')
plt.ylabel('Tip ($)')
plt.show()Encoding a Third Variable with hue
The hue parameter assigns a colour to each point based on a third variable — either categorical or numeric. When hue is a categorical column (like 'smoker'), Seaborn uses distinct colours. When hue is a numeric column, it uses a sequential colour map. This lets you see three variables at once on a 2D plot without adding a third axis.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
sns.scatterplot(
data=tips,
x='total_bill',
y='tip',
hue='time', # Lunch vs Dinner
style='smoker', # marker shape
palette='deep'
)
plt.title('Tip vs Bill by Time and Smoker Status')
plt.legend(bbox_to_anchor=(1, 1))
plt.tight_layout()
plt.show()Encoding Size and Style
Beyond hue, Seaborn scatter plots support size (point area encodes a numeric variable) and style (marker shape encodes a categorical variable). Using all three semantics simultaneously can reveal complex multi-variable patterns but risks overloading the viewer. A practical guideline is to use hue first (most salient), style second, and size only when necessary.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
sns.scatterplot(
data=tips,
x='total_bill',
y='tip',
hue='day',
size='size', # party size controls marker area
sizes=(20, 200), # min and max dot area
alpha=0.7
)
plt.title('Tip vs Bill — Colour=Day, Size=Party Size')
plt.legend(bbox_to_anchor=(1, 1))
plt.tight_layout()
plt.show()Adding a Regression Line with regplot
sns.regplot() draws a scatter plot and overlays a linear regression line with a shaded 95% confidence band. This is useful for visualising the strength and direction of the linear relationship between two variables. The confidence band narrows in regions with dense data and widens at the edges where fewer points inform the fit.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
sns.regplot(
data=tips,
x='total_bill',
y='tip',
scatter_kws={'alpha': 0.4},
line_kws={'color': 'red'}
)
plt.title('Tip vs Bill with Regression Line')
plt.show()lmplot: Faceted Regression Plots
sns.lmplot() is the figure-level version of regplot. It supports hue (separate regression lines per group in the same panel) and col/row (separate panels). This is powerful for testing whether the relationship between two variables differs across subgroups — for example, whether the bill-to-tip relationship is steeper for dinner than for lunch.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
# Separate regression line per time (Lunch / Dinner)
sns.lmplot(
data=tips,
x='total_bill',
y='tip',
hue='time',
scatter_kws={'alpha': 0.4},
height=5
)
plt.suptitle('Regression by Meal Time', y=1.02)
plt.show()Introduction to Pair Plots
sns.pairplot(df) creates a grid of scatter plots for every pair of numeric columns in the DataFrame, with distribution plots on the diagonal. This is the fastest way to visualise all pairwise relationships in a dataset in one call. For a DataFrame with n numeric columns, pairplot creates an n×n grid. It is most practical when n is between 3 and 8 — larger grids become too small to read.
import seaborn as sns
import matplotlib.pyplot as plt
iris = sns.load_dataset('iris')
# Create pairplot coloured by species
sns.pairplot(iris, hue='species', diag_kind='kde', plot_kws={'alpha': 0.5})
plt.suptitle('Iris Dataset Pairplot', y=1.02)
plt.show()Customising the Diagonal of pairplot
The diag_kind parameter controls what appears on the diagonal of the pairplot grid. Use 'hist' for histograms or 'kde' for kernel density estimates. The diagonal shows the univariate distribution of each variable, while the off-diagonal panels show pairwise scatter plots. You can also use corner=True to show only the lower triangle, halving the number of panels and reducing redundancy.
import seaborn as sns
import matplotlib.pyplot as plt
iris = sns.load_dataset('iris')
# Show only lower triangle
sns.pairplot(
iris,
hue='species',
diag_kind='hist',
corner=True
)
plt.show()PairGrid for Full Customisation
sns.PairGrid gives you full control over which plot type appears in each section of the matrix. Use g.map_upper(), g.map_lower(), and g.map_diag() to assign different functions (like regplot on the upper triangle and kdeplot on the lower). This produces publication-quality plots where each panel conveys distinct information about the variable pair.
import seaborn as sns
import matplotlib.pyplot as plt
iris = sns.load_dataset('iris')
numeric_iris = iris.drop(columns=['species'])
g = sns.PairGrid(numeric_iris)
g.map_upper(sns.scatterplot, alpha=0.3)
g.map_lower(sns.kdeplot, fill=True)
g.map_diag(sns.histplot, kde=True)
plt.suptitle('Custom PairGrid', y=1.02)
plt.show()Identifying Correlation in Scatter Plots
When reading a scatter plot, estimate the direction (does the cloud slope up or down?), strength (how tightly clustered are the points around the trend?), and form (linear or curved?) of the relationship. Correlation coefficient r ranges from -1 (perfect negative) to +1 (perfect positive), with 0 meaning no linear relationship. Remember: correlation is not causation, and a non-linear relationship can have r ≈ 0.
import seaborn as sns
import matplotlib.pyplot as plt
import numpy as np
np.random.seed(0)
x = np.linspace(0, 10, 100)
fig, axes = plt.subplots(1, 3, figsize=(12, 4))
# Strong positive correlation
axes[0].scatter(x, x + np.random.normal(0, 0.5, 100))
axes[0].set_title('Strong Positive (r≈0.99)')
# Weak correlation
axes[1].scatter(x, np.random.normal(5, 3, 100))
axes[1].set_title('No Correlation (r≈0)')
# Non-linear (quadratic) — r can be low
axes[2].scatter(x, (x-5)**2 + np.random.normal(0, 1, 100))
axes[2].set_title('Non-linear (r low despite pattern)')
plt.tight_layout()
plt.show()jointplot for Bivariate with Marginals
sns.jointplot() combines a central bivariate plot with marginal univariate plots on the top and right edges. Pass kind='scatter', 'hex', 'kde', or 'reg'. The hex kind is useful when you have thousands of overlapping points — it bins them into hexagons and uses colour to show density, solving the overplotting problem that makes scatter plots unreadable for large datasets.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
# KDE joint plot with marginal distributions
sns.jointplot(
data=tips,
x='total_bill',
y='tip',
kind='reg',
marginal_kws={'bins': 20}
)
plt.suptitle('Joint Distribution of Bill and Tip', y=1.02)
plt.show()Quick Check
Test your understanding of Seaborn scatter and pair plots from this lesson.
Lesson Recap
In this lesson you learned: sns.scatterplot encodes up to four variables using x, y, hue, size, and style, sns.regplot/lmplot overlays a regression line to quantify linear relationships, and sns.pairplot produces an all-pairs grid ideal for exploratory analysis of multi-variable datasets. Next up we explore heatmaps for visualising correlation matrices.
よくある質問
「散布図とペアプロット」レッスンは無料ですか?
はい。「散布図とペアプロット」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Pandas & NumPy Academyコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Pandas & NumPy Academyコースには全4レッスンが含まれています。
「散布図とペアプロット」で何を学びますか?
sns.scatterplotで第3の変数に応じて色分けした散布図を作成し、sns.pairplotですべてのペア間の関係を可視化します。 ブラウザで直接実行するハンズオンコードでPandas & NumPy Academyを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。
Pandas & NumPy Academyを始めるのに経験は必要ですか?
事前経験は必要ありません。CoddyKitのPandas & NumPy Academyは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン3/4です。
「散布図とペアプロット」レッスンにはどのくらい時間がかかりますか?
ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。
このPandas & NumPy Academyレッスンでコードを書いて実行できますか?
はい。すべてのPandas & NumPy Academyレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。