0Pricing
Machine Learning Academy · レッスン

顧客セグメンテーションのためのクラスタリング:一連の実例

Eコマースのデータセットを前処理し、購入金額と購入頻度で顧客をクラスタリングして、各セグメントの特徴を分析し、ビジネス上の洞察を導きます。

「顧客セグメンテーションのためのクラスタリング:一連の実例」はCoddyKit上の無料Machine Learning Academyレッスンです。 これはレッスン4/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはMachine Learning Academy学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Machine Learning Academyコースには全4レッスンが含まれています。

このレッスンの一部はまだ翻訳されておらず、英語で表示されています。

The Business Goal: Segment Customers

Customer segmentation groups buyers by behaviour so that marketing, product, and customer-success teams can tailor their actions to each group. Typical signals include recency (days since last purchase), frequency (number of purchases), and monetary value (total spend) — the RFM framework. Clustering discovers these segments from data without needing predefined categories.

Loading and Inspecting the Dataset

We use the classic Online Retail dataset (UCI ML Repository). It contains ~500k transactions with invoice date, customer ID, quantity, and unit price. Our first task is to load the data, drop rows with missing customer IDs, filter out returns (negative quantity), and compute the RFM features for each customer.

import pandas as pd

df = pd.read_csv('online_retail.csv', encoding='latin1')

# Drop missing customers and returns
df = df.dropna(subset=['CustomerID'])
df = df[df['Quantity'] > 0]
df['Revenue'] = df['Quantity'] * df['UnitPrice']
df['InvoiceDate'] = pd.to_datetime(df['InvoiceDate'])

print(df.shape)
print(df.dtypes)

Engineering RFM Features

Recency: days since the customer's last purchase (smaller = more recent = better). Frequency: number of unique invoices. Monetary: total revenue generated. We compute these relative to a snapshot date (one day after the last transaction in the dataset) so recency increases with inactivity.

snapshot_date = df['InvoiceDate'].max() + pd.Timedelta(days=1)

rfm = df.groupby('CustomerID').agg(
    Recency=('InvoiceDate', lambda x: (snapshot_date - x.max()).days),
    Frequency=('InvoiceNo', 'nunique'),
    Monetary=('Revenue', 'sum')
).reset_index()

print(rfm.describe())

Treating Outliers and Skewness

RFM features are often highly right-skewed: a handful of VIP customers dominate the monetary axis. Before scaling, apply a log transform (np.log1p) to compress the long tail. Clip extreme outliers beyond the 99th percentile to prevent a single whale customer from distorting all centroids.

import numpy as np

for col in ['Recency', 'Frequency', 'Monetary']:
    cap = rfm[col].quantile(0.99)
    rfm[col] = rfm[col].clip(upper=cap)
    rfm[col + '_log'] = np.log1p(rfm[col])

print(rfm[['Recency_log', 'Frequency_log', 'Monetary_log']].describe())

Scaling Features for K-Means

K-Means uses Euclidean distance, so features must be on the same scale. After log-transforming, apply StandardScaler to centre each feature at zero with unit variance. Always fit the scaler on training data only — here the full RFM table since there is no separate test set for unsupervised learning.

from sklearn.preprocessing import StandardScaler

features = ['Recency_log', 'Frequency_log', 'Monetary_log']
X = rfm[features].values

scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
print('Mean after scaling:', X_scaled.mean(axis=0).round(4))
print('Std after scaling:', X_scaled.std(axis=0).round(4))

Selecting k with Elbow and Silhouette

Run the elbow and silhouette diagnostics on the RFM dataset to select k. For a typical e-commerce dataset you might see the elbow around k=4 or k=5, which corresponds to intuitive segments: champions, loyal customers, at-risk customers, and churned customers.

from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score

results = []
for k in range(2, 9):
    km = KMeans(n_clusters=k, n_init=10, random_state=42)
    labels = km.fit_predict(X_scaled)
    results.append({'k': k, 'inertia': km.inertia_,
                    'silhouette': silhouette_score(X_scaled, labels)})

import pandas as pd
print(pd.DataFrame(results))

Fitting the Final Clustering Model

After selecting k, fit the final K-Means model and add the cluster labels back to the RFM DataFrame. This makes it easy to compute segment profiles and build customer-facing reports. The fit_predict method fits and returns labels in one call.

from sklearn.cluster import KMeans

k = 4
km = KMeans(n_clusters=k, n_init=20, random_state=42)
rfm['Segment'] = km.fit_predict(X_scaled)

print('Cluster sizes:')
print(rfm['Segment'].value_counts())

Profiling Each Segment

Compute the mean of the original (untransformed) RFM values for each cluster. This gives interpretable business profiles: Champions have low recency, high frequency, high monetary; Churned have high recency, low frequency, low monetary. Naming segments based on their profiles makes reports actionable.

profile = rfm.groupby('Segment')[['Recency', 'Frequency', 'Monetary']].mean()
print(profile.round(1))

# Optional: label segments by profile
segment_names = {
    0: 'Champions',
    1: 'At-Risk',
    2: 'Loyal',
    3: 'Churned'
}
rfm['SegmentName'] = rfm['Segment'].map(segment_names)
print(rfm['SegmentName'].value_counts())

Visualising Segments with Scatter Plots

Plot Frequency vs Monetary with colour coding for each segment. Add recency as point size to encode the third dimension visually. This chart is the deliverable that a marketing team can use to identify which customers to target for reactivation campaigns vs upselling campaigns.

import matplotlib.pyplot as plt

plt.figure(figsize=(8, 5))
for seg in rfm['Segment'].unique():
    mask = rfm['Segment'] == seg
    plt.scatter(rfm.loc[mask, 'Frequency'],
                rfm.loc[mask, 'Monetary'],
                s=rfm.loc[mask, 'Recency'] + 5,
                label=f'Segment {seg}', alpha=0.5)
plt.xlabel('Frequency')
plt.ylabel('Monetary')
plt.legend()
plt.title('RFM Customer Segments')
plt.show()

Assigning New Customers to Segments

After deploying the model, new customers get assigned by passing their scaled RFM vector through the same scaler and then calling km.predict. Never refit the scaler on new data — use the scaler fitted on the training RFM table to avoid shifting the feature space. The centroid positions remain fixed after fitting.

import numpy as np

# Simulate a new customer: recency=10, frequency=15, monetary=600
new_customer = np.array([[10, 15, 600]])
new_log = np.log1p(new_customer)
new_scaled = scaler.transform(new_log)

segment = km.predict(new_scaled)[0]
print('New customer segment:', segment)

Business Insights and Next Steps

Clustering is a starting point, not an end. After profiling segments, the team should design targeted actions: send re-engagement emails to At-Risk customers, offer loyalty rewards to Champions, present upsell offers to Loyal customers. Track conversion rates per segment to measure the ROI of segmentation. Periodically retrain the model as customer behaviour evolves over time.

Quick Check

Test your understanding of customer segmentation with clustering from this lesson.

Lesson Recap

In this lesson you learned: RFM (Recency, Frequency, Monetary) features are the standard building blocks for customer segmentation, log transformation and StandardScaler make skewed RFM features suitable for K-Means, and segment profiling translates cluster numbers into actionable business labels like Champions and At-Risk. Next up we explore PCA — a technique for reducing high-dimensional data to its most informative components.

よくある質問

「顧客セグメンテーションのためのクラスタリング:一連の実例」レッスンは無料ですか?

はい。「顧客セグメンテーションのためのクラスタリング:一連の実例」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Machine Learning Academyコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Machine Learning Academyコースには全4レッスンが含まれています。

「顧客セグメンテーションのためのクラスタリング:一連の実例」で何を学びますか?

Eコマースのデータセットを前処理し、購入金額と購入頻度で顧客をクラスタリングして、各セグメントの特徴を分析し、ビジネス上の洞察を導きます。 ブラウザで直接実行するハンズオンコードでMachine Learning Academyを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。

Machine Learning Academyを始めるのに経験は必要ですか?

事前経験は必要ありません。CoddyKitのMachine Learning Academyは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン4/4です。

「顧客セグメンテーションのためのクラスタリング:一連の実例」レッスンにはどのくらい時間がかかりますか?

ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。

このMachine Learning Academyレッスンでコードを書いて実行できますか?

はい。すべてのMachine Learning Academyレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。

このコースのすべてのレッスン

  1. K-Means:重心、割り当て、更新ステップ
  2. Kの選択:エルボー法とシルエットスコア
  3. DBSCAN:コア点、境界点、ノイズ
  4. 顧客セグメンテーションのためのクラスタリング:一連の実例
← Machine Learning Academyに戻る