0Pricing
Pandas & NumPy Academy · Урок

Двумерный анализ и анализ корреляций

Исследуйте взаимосвязи между парами столбцов с помощью диаграмм рассеяния, ящиков с усами и корреляционной тепловой карты.

«Двумерный анализ и анализ корреляций» — бесплатный урок Pandas & NumPy Academy на CoddyKit. Это урок 3 из 4. Ты можешь прочитать весь урок бесплатно ниже — а потом практиковать его прямо в браузере с встроенным редактором кода и ИИ-репетитором 24/7. Это часть пути обучения Pandas & NumPy Academy, и твой прогресс синхронизируется между веб-версией и приложением CoddyKit. Курс Pandas & NumPy Academy содержит 4 уроков всего.

Части этого урока еще не переведены и отображаются на английском.

From One Variable to Two

Bivariate analysis examines the relationship between exactly two columns at a time. After profiling each column individually (univariate analysis), you ask: how do these two variables relate? The analysis method depends on the combination of variable types: numeric vs. numeric (scatter plot + correlation), categorical vs. numeric (box/violin plot), or categorical vs. categorical (crosstab + chi-squared). Each combination requires a different technique.

import pandas as pd
import seaborn as sns

df = sns.load_dataset('titanic')

# Relationship types present in the dataset
print('Numeric columns:', df.select_dtypes('number').columns.tolist())
print('Categorical columns:', df.select_dtypes('object').columns.tolist())
print('Boolean columns:', df.select_dtypes('bool').columns.tolist())

Numeric vs. Numeric: Scatter Plot

For two numeric variables, a scatter plot is the primary bivariate tool. It shows each observation as a point at (x, y) coordinates, revealing the direction, strength, and form of the relationship. Always look at the plot before computing a correlation coefficient — a correlation of 0 can still show a strong non-linear (curved) relationship that the coefficient misses.

import seaborn as sns
import matplotlib.pyplot as plt

df = sns.load_dataset('titanic')

sns.scatterplot(data=df, x='age', y='fare', alpha=0.4)
plt.title('Age vs. Fare — Is There a Linear Relationship?')
plt.xlabel('Age')
plt.ylabel('Fare ($)')
plt.show()

Pearson Correlation Coefficient

The Pearson r quantifies the strength and direction of the linear relationship between two numeric variables. Compute it with df[['col1', 'col2']].corr() or scipy.stats.pearsonr(x, y) which also returns a p-value for significance testing. Always pair r with a scatter plot — high r can be driven by a few outliers, and the same r can describe very different scatter shapes (Anscombe's Quartet famously illustrates this).

import pandas as pd
import seaborn as sns
from scipy import stats

df = sns.load_dataset('titanic').dropna(subset=['age', 'fare'])

# Pearson correlation
r, p = stats.pearsonr(df['age'], df['fare'])
print(f'Pearson r = {r:.3f}, p-value = {p:.4f}')

# Spearman for robustness check
rho, p_sp = stats.spearmanr(df['age'], df['fare'])
print(f'Spearman rho = {rho:.3f}, p-value = {p_sp:.4f}')

Categorical vs. Numeric: Box and Violin Plots

When one variable is categorical and the other is numeric, the question is: does the numeric distribution differ across categories? Use a box plot or violin plot to compare. For example, does survival status affect the fare paid? Does the passenger class affect age? These plots immediately reveal whether category membership is associated with higher or lower values of the numeric variable.

import seaborn as sns
import matplotlib.pyplot as plt

df = sns.load_dataset('titanic')

fig, axes = plt.subplots(1, 2, figsize=(12, 5))

sns.boxplot(data=df, x='class', y='fare',
            order=['First', 'Second', 'Third'], ax=axes[0])
axes[0].set_title('Fare by Passenger Class')

sns.violinplot(data=df, x='survived', y='age',
               inner='quartile', ax=axes[1])
axes[1].set_title('Age Distribution by Survival')

plt.tight_layout()
plt.show()

GroupBy Statistics for Categorical vs. Numeric

Complement the visual comparison with numeric group statistics using groupby. Compute the mean, median, and count for each category to quantify the differences seen in plots. Look for categories where the mean and median diverge (indicating outliers within a group), and check whether group sizes are balanced — very small groups (<5 observations) make comparisons unreliable.

import pandas as pd
import seaborn as sns

df = sns.load_dataset('titanic')

stats = df.groupby('class')['fare'].agg(['mean', 'median', 'std', 'count'])
stats.columns = ['Mean Fare', 'Median Fare', 'Std Fare', 'Count']
print(stats.round(2))

print('\nSurvival rate by class:')
print(df.groupby('class')['survived'].mean().round(3))

Categorical vs. Categorical: Crosstabs

For two categorical variables, use pd.crosstab(df['var1'], df['var2']) to create a frequency table showing how many observations fall in each combination of categories. Normalise by row or column with normalize='index' or normalize='columns' to get proportions instead of counts. This is the foundation for studying whether two categorical variables are independent or associated.

import pandas as pd
import seaborn as sns

df = sns.load_dataset('titanic')

# Frequency table
counts = pd.crosstab(df['class'], df['survived'])
counts.columns = ['Died', 'Survived']
print('Counts:')
print(counts)

# Row-normalised proportions (survival rate per class)
props = pd.crosstab(df['class'], df['survived'], normalize='index')
props.columns = ['Died', 'Survived']
print('\nSurvival Rate per Class:')
print(props.round(3))

Visualising Crosstabs as Heatmaps

Turn a crosstab into a heatmap for instant visual comparison of cell frequencies. This is more scannable than a printed table when there are many category combinations. Annotate each cell with its value using annot=True and choose a sequential colour map (like 'YlOrRd') since all values are non-negative. Heatmaps of row-normalised proportions effectively show the conditional distribution of one category given another.

import pandas as pd
import seaborn as sns
import matplotlib.pyplot as plt

df = sns.load_dataset('titanic')
props = pd.crosstab(df['class'], df['sex'], normalize='index')

sns.heatmap(props, annot=True, fmt='.2f', cmap='YlOrRd')
plt.title('Gender Mix by Passenger Class (Row %)')
plt.xlabel('Sex')
plt.ylabel('Class')
plt.show()

Chi-Squared Test for Independence

The chi-squared test of independence checks whether two categorical variables are statistically independent. scipy.stats.chi2_contingency(crosstab) returns the chi-squared statistic, p-value, degrees of freedom, and expected frequencies. A small p-value (typically < 0.05) means there is a statistically significant association between the two variables — the distribution of one varies systematically with the other.

import pandas as pd
import seaborn as sns
from scipy.stats import chi2_contingency

df = sns.load_dataset('titanic')

# Test if passenger class and survival are independent
crosstab = pd.crosstab(df['class'], df['survived'])
chi2, p, dof, expected = chi2_contingency(crosstab)

print(f'Chi-squared: {chi2:.2f}')
print(f'P-value: {p:.6f}')
print(f'Degrees of freedom: {dof}')
print(f'\nConclusion: {"Significant association" if p < 0.05 else "No significant association"}')

Full Correlation Matrix for Numeric Columns

Rather than examining pairs one at a time, compute the full correlation matrix for all numeric columns and visualise it as a heatmap. This reveals which pairs of variables are strongly correlated in one glance. High correlations between independent features (multicollinearity) can cause problems in regression models — knowing about them early allows you to remove or combine redundant features before modelling.

import pandas as pd
import seaborn as sns
import matplotlib.pyplot as plt
import numpy as np

df = sns.load_dataset('titanic')
corr = df.select_dtypes('number').corr()
mask = np.triu(np.ones_like(corr, dtype=bool))

sns.heatmap(corr, mask=mask, annot=True, fmt='.2f',
            cmap='coolwarm', vmin=-1, vmax=1,
            linewidths=0.5, square=True)
plt.title('Correlation Matrix — Titanic Numeric Columns')
plt.tight_layout()
plt.show()

Interactions Across Subgroups

Sometimes a relationship between two variables differs dramatically across subgroups — a phenomenon known as Simpson's Paradox. For example, the correlation between age and fare might be positive for first-class passengers but negative for third-class. Always segment your bivariate analysis by key categorical variables (using hue in Seaborn or separate groupby aggregations) to check whether the relationship is consistent or reverses across subgroups.

import seaborn as sns
import matplotlib.pyplot as plt

df = sns.load_dataset('titanic')

# Scatter coloured by passenger class
sns.scatterplot(
    data=df.dropna(subset=['age', 'fare']),
    x='age',
    y='fare',
    hue='class',
    alpha=0.5,
    palette='deep'
)
plt.title('Age vs Fare — Does Class Change the Relationship?')
plt.legend(title='Class')
plt.yscale('log')  # log scale due to fare skewness
plt.show()

Documenting Bivariate Findings

After completing bivariate analysis, document key findings: which pairs of features are correlated (and how strongly), whether categorical variables show meaningful group differences in numeric outcomes, and any subgroup interactions or paradoxes found. These insights should drive feature selection decisions — highly correlated feature pairs may need one dropped, and strong categorical predictors of the target variable should be flagged as high-priority inputs for modelling.

Quick Check

Test your understanding of bivariate and correlation analysis from this lesson.

Lesson Recap

In this lesson you learned: scatter plots and Pearson r analyse numeric-numeric relationships, box/violin plots and groupby compare numeric variables across categories, and crosstabs with chi-squared tests measure association between categorical variables. Next up we cover how to compile EDA findings into a structured summary report.

Часто задаваемые вопросы

Урок «Двумерный анализ и анализ корреляций» бесплатный?

Да — полный текст урока «Двумерный анализ и анализ корреляций» бесплатно доступен здесь в веб-версии. Чтобы практиковать его интерактивно (встроенный редактор кода и ИИ-репетитор 24/7) и разблокировать остальной курс Pandas & NumPy Academy, подпишись на CoddyKit PRO. Курс Pandas & NumPy Academy содержит 4 уроков всего.

Чему я научусь в уроке «Двумерный анализ и анализ корреляций»?

Исследуйте взаимосвязи между парами столбцов с помощью диаграмм рассеяния, ящиков с усами и корреляционной тепловой карты. Ты практикуешь Pandas & NumPy Academy с помощью реального кода, который запускаешь прямо в браузере, и ИИ-репетитор 24/7 отвечает на твои вопросы во время урока.

Нужен ли мне опыт, чтобы начать Pandas & NumPy Academy?

Предыдущий опыт не требуется. Pandas & NumPy Academy на CoddyKit структурирован для всех уровней — от новичков до продвинутых, поэтому ты можешь начать отсюда или с самого начала и учиться в своем темпе. Это урок 3 из 4.

Сколько времени занимает урок «Двумерный анализ и анализ корреляций»?

Большинство уроков CoddyKit занимают около 5–10 минут. Каждый из них компактный и интерактивный, поэтому ты постоянно делаешь прогресс и продолжаешь с того же места в веб-версии и приложении.

Можно ли писать и запускать код в этом уроке Pandas & NumPy Academy?

Да. Каждый урок Pandas & NumPy Academy включает встроенный редактор кода, поэтому ты пишешь и запускаешь реальный код прямо в браузере и получаешь моментальную обратную связь от AI — локальная установка не требуется.

Все уроки этого курса

  1. Контрольный список профилирования набора данных
  2. Одномерный анализ
  3. Двумерный анализ и анализ корреляций
  4. Обобщение результатов в отчёте
← Назад к Pandas & NumPy Academy