Machine Learning Academy · Ders

Müşteri Segmentasyonu için Kümeleme: Baştan Sona Örnek

Bir e-ticaret veri kümesini ön işlemeden geçirecek, müşterileri harcama ve alışveriş sıklığına göre kümeleyecek, iş içgörüleri üretmek için her segmentin profilini çıkaracaksınız.

4. ders / 413 adım

Müşteri Segmentasyonu için Kümeleme: Baştan Sona Örnek, CoddyKit'te ücretsiz bir Machine Learning Academy dersidir. Bu, 4 dersinin 4. dersidir. Aşağıdan dersin tamamını ücretsiz okuyabilir, sonra tarayıcıda yerleşik kod editörü ve 7/24 yapay zeka koçu ile uygulamalı olarak pratik yapabilirsin. Bu, Machine Learning Academy öğrenme yolunun bir parçasıdır ve ilerlemeniz web ve CoddyKit uygulaması arasında senkronize olur. Machine Learning Academy kursu toplamda 4 dersten oluşur.

Bu dersin bazı bölümleri henüz çevrilmemiş olup İngilizce olarak gösterilmektedir.

The Business Goal: Segment Customers

Customer segmentation groups buyers by behaviour so that marketing, product, and customer-success teams can tailor their actions to each group. Typical signals include recency (days since last purchase), frequency (number of purchases), and monetary value (total spend) — the RFM framework. Clustering discovers these segments from data without needing predefined categories.

Loading and Inspecting the Dataset

We use the classic Online Retail dataset (UCI ML Repository). It contains ~500k transactions with invoice date, customer ID, quantity, and unit price. Our first task is to load the data, drop rows with missing customer IDs, filter out returns (negative quantity), and compute the RFM features for each customer.

import pandas as pd

df = pd.read_csv('online_retail.csv', encoding='latin1')

# Drop missing customers and returns
df = df.dropna(subset=['CustomerID'])
df = df[df['Quantity'] > 0]
df['Revenue'] = df['Quantity'] * df['UnitPrice']
df['InvoiceDate'] = pd.to_datetime(df['InvoiceDate'])

print(df.shape)
print(df.dtypes)

Engineering RFM Features

Recency: days since the customer's last purchase (smaller = more recent = better). Frequency: number of unique invoices. Monetary: total revenue generated. We compute these relative to a snapshot date (one day after the last transaction in the dataset) so recency increases with inactivity.

snapshot_date = df['InvoiceDate'].max() + pd.Timedelta(days=1)

rfm = df.groupby('CustomerID').agg(
    Recency=('InvoiceDate', lambda x: (snapshot_date - x.max()).days),
    Frequency=('InvoiceNo', 'nunique'),
    Monetary=('Revenue', 'sum')
).reset_index()

print(rfm.describe())

Treating Outliers and Skewness

RFM features are often highly right-skewed: a handful of VIP customers dominate the monetary axis. Before scaling, apply a log transform (np.log1p) to compress the long tail. Clip extreme outliers beyond the 99th percentile to prevent a single whale customer from distorting all centroids.

import numpy as np

for col in ['Recency', 'Frequency', 'Monetary']:
    cap = rfm[col].quantile(0.99)
    rfm[col] = rfm[col].clip(upper=cap)
    rfm[col + '_log'] = np.log1p(rfm[col])

print(rfm[['Recency_log', 'Frequency_log', 'Monetary_log']].describe())

Scaling Features for K-Means

K-Means uses Euclidean distance, so features must be on the same scale. After log-transforming, apply StandardScaler to centre each feature at zero with unit variance. Always fit the scaler on training data only — here the full RFM table since there is no separate test set for unsupervised learning.

from sklearn.preprocessing import StandardScaler

features = ['Recency_log', 'Frequency_log', 'Monetary_log']
X = rfm[features].values

scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
print('Mean after scaling:', X_scaled.mean(axis=0).round(4))
print('Std after scaling:', X_scaled.std(axis=0).round(4))

Selecting k with Elbow and Silhouette

Run the elbow and silhouette diagnostics on the RFM dataset to select k. For a typical e-commerce dataset you might see the elbow around k=4 or k=5, which corresponds to intuitive segments: champions, loyal customers, at-risk customers, and churned customers.

from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score

results = []
for k in range(2, 9):
    km = KMeans(n_clusters=k, n_init=10, random_state=42)
    labels = km.fit_predict(X_scaled)
    results.append({'k': k, 'inertia': km.inertia_,
                    'silhouette': silhouette_score(X_scaled, labels)})

import pandas as pd
print(pd.DataFrame(results))

Fitting the Final Clustering Model

After selecting k, fit the final K-Means model and add the cluster labels back to the RFM DataFrame. This makes it easy to compute segment profiles and build customer-facing reports. The fit_predict method fits and returns labels in one call.

from sklearn.cluster import KMeans

k = 4
km = KMeans(n_clusters=k, n_init=20, random_state=42)
rfm['Segment'] = km.fit_predict(X_scaled)

print('Cluster sizes:')
print(rfm['Segment'].value_counts())

Profiling Each Segment

Compute the mean of the original (untransformed) RFM values for each cluster. This gives interpretable business profiles: Champions have low recency, high frequency, high monetary; Churned have high recency, low frequency, low monetary. Naming segments based on their profiles makes reports actionable.

profile = rfm.groupby('Segment')[['Recency', 'Frequency', 'Monetary']].mean()
print(profile.round(1))

# Optional: label segments by profile
segment_names = {
    0: 'Champions',
    1: 'At-Risk',
    2: 'Loyal',
    3: 'Churned'
}
rfm['SegmentName'] = rfm['Segment'].map(segment_names)
print(rfm['SegmentName'].value_counts())

Visualising Segments with Scatter Plots

Plot Frequency vs Monetary with colour coding for each segment. Add recency as point size to encode the third dimension visually. This chart is the deliverable that a marketing team can use to identify which customers to target for reactivation campaigns vs upselling campaigns.

import matplotlib.pyplot as plt

plt.figure(figsize=(8, 5))
for seg in rfm['Segment'].unique():
    mask = rfm['Segment'] == seg
    plt.scatter(rfm.loc[mask, 'Frequency'],
                rfm.loc[mask, 'Monetary'],
                s=rfm.loc[mask, 'Recency'] + 5,
                label=f'Segment {seg}', alpha=0.5)
plt.xlabel('Frequency')
plt.ylabel('Monetary')
plt.legend()
plt.title('RFM Customer Segments')
plt.show()

Assigning New Customers to Segments

After deploying the model, new customers get assigned by passing their scaled RFM vector through the same scaler and then calling km.predict. Never refit the scaler on new data — use the scaler fitted on the training RFM table to avoid shifting the feature space. The centroid positions remain fixed after fitting.

import numpy as np

# Simulate a new customer: recency=10, frequency=15, monetary=600
new_customer = np.array([[10, 15, 600]])
new_log = np.log1p(new_customer)
new_scaled = scaler.transform(new_log)

segment = km.predict(new_scaled)[0]
print('New customer segment:', segment)

Business Insights and Next Steps

Clustering is a starting point, not an end. After profiling segments, the team should design targeted actions: send re-engagement emails to At-Risk customers, offer loyalty rewards to Champions, present upsell offers to Loyal customers. Track conversion rates per segment to measure the ROI of segmentation. Periodically retrain the model as customer behaviour evolves over time.

Quick Check

Test your understanding of customer segmentation with clustering from this lesson.

Lesson Recap

In this lesson you learned: RFM (Recency, Frequency, Monetary) features are the standard building blocks for customer segmentation, log transformation and StandardScaler make skewed RFM features suitable for K-Means, and segment profiling translates cluster numbers into actionable business labels like Champions and At-Risk. Next up we explore PCA — a technique for reducing high-dimensional data to its most informative components.

Başlamak ücretsiz

Yapay zeka eğitmeniyle Python öğren — ücretsiz

Tarayıcında gerçek kod yaz ve çalıştır, 7/24 yapay zeka eğitmeninden anında yardım al; web'de ya da uygulamada kaldığın yerden devam et.

Kurslar
30
Dersler
120

Sıkça Sorulan Sorular

“Müşteri Segmentasyonu için Kümeleme: Baştan Sona Örnek” dersi ücretsiz mi?

Evet — “Müşteri Segmentasyonu için Kümeleme: Baştan Sona Örnek” dersin tüm metni burada web'de ücretsiz olarak okunabilir. Etkileşimli olarak pratik yapmak (yerleşik kod editörü ve 7/24 yapay zeka koçu) ve Machine Learning Academy kursunun geri kalanını açmak için CoddyKit PRO'ya yükselt. Machine Learning Academy kursu toplamda 4 dersten oluşur.

“Müşteri Segmentasyonu için Kümeleme: Baştan Sona Örnek” dersinde ne öğreneceğim?

Bir e-ticaret veri kümesini ön işlemeden geçirecek, müşterileri harcama ve alışveriş sıklığına göre kümeleyecek, iş içgörüleri üretmek için her segmentin profilini çıkaracaksınız. Machine Learning Academy ile uygulamalı kodu tarayıcıda doğrudan çalıştırarak pratik yaparsın ve 7/24 yapay zeka koçu dersi çalışırken sorularını yanıtlar.

Machine Learning Academy öğrenmeye başlamak için deneyim gerekli mi?

Önceden deneyim gerekmez. CoddyKit'te Machine Learning Academy, başlangıçtan ileri seviyeye kadar yapılandırıldığı için buradan başlayabilir veya başından başlayıp kendi hızında ilerleme yapabilirsin. Bu, 4 dersinin 4. dersidir.

“Müşteri Segmentasyonu için Kümeleme: Baştan Sona Örnek” dersi ne kadar sürer?

Çoğu CoddyKit dersi yaklaşık 5–10 dakika sürer. Her biri kısa ve etkileşimli olduğu için sabit ilerleme yaparsın ve web ile uygulama arasında tam olarak bıraktığın yerden devam edebilirsin.

Bu Machine Learning Academy dersinde kod yazıp çalıştırabilir miyim?

Evet. Her Machine Learning Academy dersi yerleşik bir kod editörü içerir, bu sayede tarayıcıda gerçek kod yazıp çalıştırabilir ve anlık yapay zeka geri bildirimi alırsın — yerel kurulum gerekli değildir.

Bu kursun tüm dersleri

  1. K-Means: Ağırlık Merkezleri, Atama ve Güncelleme Adımları
  2. K Seçimi: Dirsek Yöntemi ve Siluet Skoru
  3. DBSCAN: Çekirdek Noktalar, Sınır Noktaları ve Gürültü
  4. Müşteri Segmentasyonu için Kümeleme: Baştan Sona Örnek
← Machine Learning Academy Sayfasına Dön