Streudiagramme und Pairplots
Erstellen Sie mit sns.scatterplot nach einer dritten Variable eingefärbte Streudiagramme und visualisieren Sie mit sns.pairplot alle paarweisen Beziehungen.
Streudiagramme und Pairplots ist eine kostenlose Pandas & NumPy Academy-Lektion auf CoddyKit. Dies ist Lektion 3 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des Pandas & NumPy Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der Pandas & NumPy Academy-Kurs umfasst insgesamt 4 Lektionen.
Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.
Visualising Relationships Between Variables
Distribution plots show one variable at a time. When you want to understand the relationship between two numeric variables, you need scatter plots and related visualisations. These help you detect correlation (do both variables increase together?), clusters (are there subgroups in the data?), and outliers (are there unusual points far from the main cloud?). Seaborn makes all of these easy with scatterplot and pairplot.
import seaborn as sns
import matplotlib.pyplot as plt
# Load the classic iris dataset
iris = sns.load_dataset('iris')
print(iris.head())
print('\nShape:', iris.shape)Basic Scatter Plot with sns.scatterplot
sns.scatterplot(data, x, y) plots each observation as a point at its (x, y) coordinates. Unlike Matplotlib's plt.scatter, Seaborn's version automatically links to a DataFrame and integrates with hue, size, and style semantics. The result is informative multi-variable plots with a single function call and an automatic legend.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
sns.scatterplot(data=tips, x='total_bill', y='tip')
plt.title('Tip vs Total Bill')
plt.xlabel('Total Bill ($)')
plt.ylabel('Tip ($)')
plt.show()Encoding a Third Variable with hue
The hue parameter assigns a colour to each point based on a third variable — either categorical or numeric. When hue is a categorical column (like 'smoker'), Seaborn uses distinct colours. When hue is a numeric column, it uses a sequential colour map. This lets you see three variables at once on a 2D plot without adding a third axis.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
sns.scatterplot(
data=tips,
x='total_bill',
y='tip',
hue='time', # Lunch vs Dinner
style='smoker', # marker shape
palette='deep'
)
plt.title('Tip vs Bill by Time and Smoker Status')
plt.legend(bbox_to_anchor=(1, 1))
plt.tight_layout()
plt.show()Encoding Size and Style
Beyond hue, Seaborn scatter plots support size (point area encodes a numeric variable) and style (marker shape encodes a categorical variable). Using all three semantics simultaneously can reveal complex multi-variable patterns but risks overloading the viewer. A practical guideline is to use hue first (most salient), style second, and size only when necessary.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
sns.scatterplot(
data=tips,
x='total_bill',
y='tip',
hue='day',
size='size', # party size controls marker area
sizes=(20, 200), # min and max dot area
alpha=0.7
)
plt.title('Tip vs Bill — Colour=Day, Size=Party Size')
plt.legend(bbox_to_anchor=(1, 1))
plt.tight_layout()
plt.show()Adding a Regression Line with regplot
sns.regplot() draws a scatter plot and overlays a linear regression line with a shaded 95% confidence band. This is useful for visualising the strength and direction of the linear relationship between two variables. The confidence band narrows in regions with dense data and widens at the edges where fewer points inform the fit.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
sns.regplot(
data=tips,
x='total_bill',
y='tip',
scatter_kws={'alpha': 0.4},
line_kws={'color': 'red'}
)
plt.title('Tip vs Bill with Regression Line')
plt.show()lmplot: Faceted Regression Plots
sns.lmplot() is the figure-level version of regplot. It supports hue (separate regression lines per group in the same panel) and col/row (separate panels). This is powerful for testing whether the relationship between two variables differs across subgroups — for example, whether the bill-to-tip relationship is steeper for dinner than for lunch.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
# Separate regression line per time (Lunch / Dinner)
sns.lmplot(
data=tips,
x='total_bill',
y='tip',
hue='time',
scatter_kws={'alpha': 0.4},
height=5
)
plt.suptitle('Regression by Meal Time', y=1.02)
plt.show()Introduction to Pair Plots
sns.pairplot(df) creates a grid of scatter plots for every pair of numeric columns in the DataFrame, with distribution plots on the diagonal. This is the fastest way to visualise all pairwise relationships in a dataset in one call. For a DataFrame with n numeric columns, pairplot creates an n×n grid. It is most practical when n is between 3 and 8 — larger grids become too small to read.
import seaborn as sns
import matplotlib.pyplot as plt
iris = sns.load_dataset('iris')
# Create pairplot coloured by species
sns.pairplot(iris, hue='species', diag_kind='kde', plot_kws={'alpha': 0.5})
plt.suptitle('Iris Dataset Pairplot', y=1.02)
plt.show()Customising the Diagonal of pairplot
The diag_kind parameter controls what appears on the diagonal of the pairplot grid. Use 'hist' for histograms or 'kde' for kernel density estimates. The diagonal shows the univariate distribution of each variable, while the off-diagonal panels show pairwise scatter plots. You can also use corner=True to show only the lower triangle, halving the number of panels and reducing redundancy.
import seaborn as sns
import matplotlib.pyplot as plt
iris = sns.load_dataset('iris')
# Show only lower triangle
sns.pairplot(
iris,
hue='species',
diag_kind='hist',
corner=True
)
plt.show()PairGrid for Full Customisation
sns.PairGrid gives you full control over which plot type appears in each section of the matrix. Use g.map_upper(), g.map_lower(), and g.map_diag() to assign different functions (like regplot on the upper triangle and kdeplot on the lower). This produces publication-quality plots where each panel conveys distinct information about the variable pair.
import seaborn as sns
import matplotlib.pyplot as plt
iris = sns.load_dataset('iris')
numeric_iris = iris.drop(columns=['species'])
g = sns.PairGrid(numeric_iris)
g.map_upper(sns.scatterplot, alpha=0.3)
g.map_lower(sns.kdeplot, fill=True)
g.map_diag(sns.histplot, kde=True)
plt.suptitle('Custom PairGrid', y=1.02)
plt.show()Identifying Correlation in Scatter Plots
When reading a scatter plot, estimate the direction (does the cloud slope up or down?), strength (how tightly clustered are the points around the trend?), and form (linear or curved?) of the relationship. Correlation coefficient r ranges from -1 (perfect negative) to +1 (perfect positive), with 0 meaning no linear relationship. Remember: correlation is not causation, and a non-linear relationship can have r ≈ 0.
import seaborn as sns
import matplotlib.pyplot as plt
import numpy as np
np.random.seed(0)
x = np.linspace(0, 10, 100)
fig, axes = plt.subplots(1, 3, figsize=(12, 4))
# Strong positive correlation
axes[0].scatter(x, x + np.random.normal(0, 0.5, 100))
axes[0].set_title('Strong Positive (r≈0.99)')
# Weak correlation
axes[1].scatter(x, np.random.normal(5, 3, 100))
axes[1].set_title('No Correlation (r≈0)')
# Non-linear (quadratic) — r can be low
axes[2].scatter(x, (x-5)**2 + np.random.normal(0, 1, 100))
axes[2].set_title('Non-linear (r low despite pattern)')
plt.tight_layout()
plt.show()jointplot for Bivariate with Marginals
sns.jointplot() combines a central bivariate plot with marginal univariate plots on the top and right edges. Pass kind='scatter', 'hex', 'kde', or 'reg'. The hex kind is useful when you have thousands of overlapping points — it bins them into hexagons and uses colour to show density, solving the overplotting problem that makes scatter plots unreadable for large datasets.
import seaborn as sns
import matplotlib.pyplot as plt
tips = sns.load_dataset('tips')
# KDE joint plot with marginal distributions
sns.jointplot(
data=tips,
x='total_bill',
y='tip',
kind='reg',
marginal_kws={'bins': 20}
)
plt.suptitle('Joint Distribution of Bill and Tip', y=1.02)
plt.show()Quick Check
Test your understanding of Seaborn scatter and pair plots from this lesson.
Lesson Recap
In this lesson you learned: sns.scatterplot encodes up to four variables using x, y, hue, size, and style, sns.regplot/lmplot overlays a regression line to quantify linear relationships, and sns.pairplot produces an all-pairs grid ideal for exploratory analysis of multi-variable datasets. Next up we explore heatmaps for visualising correlation matrices.
Häufig gestellte Fragen
Ist die Lektion „Streudiagramme und Pairplots“ kostenlos?
Ja — der vollständige Text von „Streudiagramme und Pairplots“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des Pandas & NumPy Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der Pandas & NumPy Academy-Kurs umfasst insgesamt 4 Lektionen.
Was lerne ich in „Streudiagramme und Pairplots“?
Erstellen Sie mit sns.scatterplot nach einer dritten Variable eingefärbte Streudiagramme und visualisieren Sie mit sns.pairplot alle paarweisen Beziehungen. Du übst Pandas & NumPy Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.
Brauche ich Erfahrung, um Pandas & NumPy Academy zu starten?
Keine Vorkenntnisse erforderlich. Pandas & NumPy Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 3 von 4.
Wie lange dauert die Lektion „Streudiagramme und Pairplots“?
Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.
Kann ich in dieser Pandas & NumPy Academy-Lektion Code schreiben und ausführen?
Ja. Jede Pandas & NumPy Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.
Alle Lektionen in diesem Kurs
- Verteilungsdiagramme: histplot und kdeplot
- Kategoriale Diagramme: boxplot, barplot, violinplot
- Streudiagramme und Pairplots
- Heatmaps für Korrelationsmatrizen