用于客户细分的聚类:端到端示例
您将预处理电子商务数据集,依据消费金额和频率对客户进行聚类,并分析每个细分群体以获得业务洞察。
用于客户细分的聚类:端到端示例 是 CoddyKit 上的免费 Machine Learning Academy 课时。 这是第 4 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Machine Learning Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Machine Learning Academy 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
The Business Goal: Segment Customers
Customer segmentation groups buyers by behaviour so that marketing, product, and customer-success teams can tailor their actions to each group. Typical signals include recency (days since last purchase), frequency (number of purchases), and monetary value (total spend) — the RFM framework. Clustering discovers these segments from data without needing predefined categories.
Loading and Inspecting the Dataset
We use the classic Online Retail dataset (UCI ML Repository). It contains ~500k transactions with invoice date, customer ID, quantity, and unit price. Our first task is to load the data, drop rows with missing customer IDs, filter out returns (negative quantity), and compute the RFM features for each customer.
import pandas as pd
df = pd.read_csv('online_retail.csv', encoding='latin1')
# Drop missing customers and returns
df = df.dropna(subset=['CustomerID'])
df = df[df['Quantity'] > 0]
df['Revenue'] = df['Quantity'] * df['UnitPrice']
df['InvoiceDate'] = pd.to_datetime(df['InvoiceDate'])
print(df.shape)
print(df.dtypes)Engineering RFM Features
Recency: days since the customer's last purchase (smaller = more recent = better). Frequency: number of unique invoices. Monetary: total revenue generated. We compute these relative to a snapshot date (one day after the last transaction in the dataset) so recency increases with inactivity.
snapshot_date = df['InvoiceDate'].max() + pd.Timedelta(days=1)
rfm = df.groupby('CustomerID').agg(
Recency=('InvoiceDate', lambda x: (snapshot_date - x.max()).days),
Frequency=('InvoiceNo', 'nunique'),
Monetary=('Revenue', 'sum')
).reset_index()
print(rfm.describe())Treating Outliers and Skewness
RFM features are often highly right-skewed: a handful of VIP customers dominate the monetary axis. Before scaling, apply a log transform (np.log1p) to compress the long tail. Clip extreme outliers beyond the 99th percentile to prevent a single whale customer from distorting all centroids.
import numpy as np
for col in ['Recency', 'Frequency', 'Monetary']:
cap = rfm[col].quantile(0.99)
rfm[col] = rfm[col].clip(upper=cap)
rfm[col + '_log'] = np.log1p(rfm[col])
print(rfm[['Recency_log', 'Frequency_log', 'Monetary_log']].describe())Scaling Features for K-Means
K-Means uses Euclidean distance, so features must be on the same scale. After log-transforming, apply StandardScaler to centre each feature at zero with unit variance. Always fit the scaler on training data only — here the full RFM table since there is no separate test set for unsupervised learning.
from sklearn.preprocessing import StandardScaler
features = ['Recency_log', 'Frequency_log', 'Monetary_log']
X = rfm[features].values
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
print('Mean after scaling:', X_scaled.mean(axis=0).round(4))
print('Std after scaling:', X_scaled.std(axis=0).round(4))Selecting k with Elbow and Silhouette
Run the elbow and silhouette diagnostics on the RFM dataset to select k. For a typical e-commerce dataset you might see the elbow around k=4 or k=5, which corresponds to intuitive segments: champions, loyal customers, at-risk customers, and churned customers.
from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score
results = []
for k in range(2, 9):
km = KMeans(n_clusters=k, n_init=10, random_state=42)
labels = km.fit_predict(X_scaled)
results.append({'k': k, 'inertia': km.inertia_,
'silhouette': silhouette_score(X_scaled, labels)})
import pandas as pd
print(pd.DataFrame(results))Fitting the Final Clustering Model
After selecting k, fit the final K-Means model and add the cluster labels back to the RFM DataFrame. This makes it easy to compute segment profiles and build customer-facing reports. The fit_predict method fits and returns labels in one call.
from sklearn.cluster import KMeans
k = 4
km = KMeans(n_clusters=k, n_init=20, random_state=42)
rfm['Segment'] = km.fit_predict(X_scaled)
print('Cluster sizes:')
print(rfm['Segment'].value_counts())Profiling Each Segment
Compute the mean of the original (untransformed) RFM values for each cluster. This gives interpretable business profiles: Champions have low recency, high frequency, high monetary; Churned have high recency, low frequency, low monetary. Naming segments based on their profiles makes reports actionable.
profile = rfm.groupby('Segment')[['Recency', 'Frequency', 'Monetary']].mean()
print(profile.round(1))
# Optional: label segments by profile
segment_names = {
0: 'Champions',
1: 'At-Risk',
2: 'Loyal',
3: 'Churned'
}
rfm['SegmentName'] = rfm['Segment'].map(segment_names)
print(rfm['SegmentName'].value_counts())Visualising Segments with Scatter Plots
Plot Frequency vs Monetary with colour coding for each segment. Add recency as point size to encode the third dimension visually. This chart is the deliverable that a marketing team can use to identify which customers to target for reactivation campaigns vs upselling campaigns.
import matplotlib.pyplot as plt
plt.figure(figsize=(8, 5))
for seg in rfm['Segment'].unique():
mask = rfm['Segment'] == seg
plt.scatter(rfm.loc[mask, 'Frequency'],
rfm.loc[mask, 'Monetary'],
s=rfm.loc[mask, 'Recency'] + 5,
label=f'Segment {seg}', alpha=0.5)
plt.xlabel('Frequency')
plt.ylabel('Monetary')
plt.legend()
plt.title('RFM Customer Segments')
plt.show()Assigning New Customers to Segments
After deploying the model, new customers get assigned by passing their scaled RFM vector through the same scaler and then calling km.predict. Never refit the scaler on new data — use the scaler fitted on the training RFM table to avoid shifting the feature space. The centroid positions remain fixed after fitting.
import numpy as np
# Simulate a new customer: recency=10, frequency=15, monetary=600
new_customer = np.array([[10, 15, 600]])
new_log = np.log1p(new_customer)
new_scaled = scaler.transform(new_log)
segment = km.predict(new_scaled)[0]
print('New customer segment:', segment)Business Insights and Next Steps
Clustering is a starting point, not an end. After profiling segments, the team should design targeted actions: send re-engagement emails to At-Risk customers, offer loyalty rewards to Champions, present upsell offers to Loyal customers. Track conversion rates per segment to measure the ROI of segmentation. Periodically retrain the model as customer behaviour evolves over time.
Quick Check
Test your understanding of customer segmentation with clustering from this lesson.
Lesson Recap
In this lesson you learned: RFM (Recency, Frequency, Monetary) features are the standard building blocks for customer segmentation, log transformation and StandardScaler make skewed RFM features suitable for K-Means, and segment profiling translates cluster numbers into actionable business labels like Champions and At-Risk. Next up we explore PCA — a technique for reducing high-dimensional data to its most informative components.
常见问题解答
「用于客户细分的聚类:端到端示例」课时是免费的吗?
是的 — 「用于客户细分的聚类:端到端示例」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Machine Learning Academy 课程的其余内容,请升级到 CoddyKit PRO。 Machine Learning Academy 课程共包含 4 节课。
「用于客户细分的聚类:端到端示例」这节课中我会学到什么?
您将预处理电子商务数据集,依据消费金额和频率对客户进行聚类,并分析每个细分群体以获得业务洞察。 你通过在浏览器中直接运行的动手代码来练习 Machine Learning Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Machine Learning Academy 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Machine Learning Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 4 节课,共 4 节。
「用于客户细分的聚类:端到端示例」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Machine Learning Academy 课中编写并运行代码吗?
能。每节 Machine Learning Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- K-Means:质心、分配与更新步骤
- 选择 K:肘部法与轮廓系数
- DBSCAN:核心点、边界点与噪声点
- 用于客户细分的聚类:端到端示例