0Pricing
Machine Learning Academy · Lekcja

Wybór K: metoda łokcia i współczynnik sylwetki

Uczą się Państwo rysować zależność inercji od k (metoda łokcia) oraz obliczać współczynniki sylwetki, aby wybrać liczbę klastrów tworzących zwarte i dobrze odseparowane grupy.

Wybór K: metoda łokcia i współczynnik sylwetki to bezpłatna lekcja Machine Learning Academy na CoddyKit. To lekcja 2 z 4. Możesz przeczytać całą lekcję poniżej za darmo — a potem ćwiczyć ją interaktywnie w przeglądarce z wbudowanym edytorem kodu i tutorem AI dostępnym 24/7. To część ścieżki edukacyjnej Machine Learning Academy, a Twój postęp synchronizuje się między webem a aplikacją CoddyKit. Kurs Machine Learning Academy zawiera 4 lekcji w sumie.

Części tej lekcji nie zostały jeszcze przetłumaczone i są wyświetlane po angielsku.

Why Choosing k Matters

K-Means requires you to specify k — the number of clusters — before training. Too few clusters and you lump distinct groups together; too many and you split natural groups artificially. There is no universally correct k, but two diagnostic tools — the elbow method and the silhouette score — give principled guidance.

Inertia Decreases as k Grows

As you increase k, inertia always decreases because points are assigned to closer centroids. At k=n (one cluster per point), inertia is zero. This means you cannot simply minimise inertia — you need to find where additional clusters stop providing meaningful reductions. That point of diminishing returns is the elbow.

from sklearn.cluster import KMeans
import numpy as np

X = np.random.randn(200, 2)
inertias = []

for k in range(1, 11):
    km = KMeans(n_clusters=k, random_state=42, n_init=10)
    km.fit(X)
    inertias.append(km.inertia_)

print('Inertia per k:')
for k, inr in enumerate(inertias, start=1):
    print(f'  k={k}: {inr:.1f}')

The Elbow Method Explained

Plot inertia on the y-axis against k on the x-axis. The curve typically drops steeply for the first few k values then flattens. The elbow — the kink where the rate of decrease sharply slows — is your estimate of the true cluster count. If the true k is 3, the drop from k=1 to k=3 is large, but from k=3 to k=4 is much smaller.

import matplotlib.pyplot as plt
from sklearn.cluster import KMeans
from sklearn.datasets import make_blobs

X, _ = make_blobs(n_samples=300, centers=4, cluster_std=0.7, random_state=0)

inertias = []
for k in range(1, 11):
    km = KMeans(n_clusters=k, random_state=0, n_init=10)
    km.fit(X)
    inertias.append(km.inertia_)

plt.plot(range(1, 11), inertias, marker='o')
plt.xlabel('Number of clusters k')
plt.ylabel('Inertia')
plt.title('Elbow Method')
plt.axvline(x=4, color='red', linestyle='--', label='True k=4')
plt.legend()
plt.show()

Limitations of the Elbow Method

The elbow method works well when clusters are clearly separated, but real-world data often produces a smooth curve with no obvious kink. In such cases the elbow is ambiguous and different people may pick different k. That is where the silhouette score provides a more objective, mathematically grounded alternative.

Silhouette Score: The Formula

For each point i, compute two values: a(i) = mean distance to other points in the same cluster (cohesion), and b(i) = mean distance to the nearest different cluster (separation). The silhouette for point i is s(i) = (b(i) - a(i)) / max(a(i), b(i)). Values range from −1 (wrong cluster) through 0 (on border) to +1 (tight, well-separated cluster).

Computing Silhouette Score in sklearn

sklearn.metrics.silhouette_score returns the mean silhouette over all points. A score above 0.5 typically indicates reasonable clustering; above 0.7 is strong. Because you cannot compute the silhouette for k=1 (no second cluster), sweep k from 2 to some maximum and pick the k with the highest mean score.

from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score
from sklearn.datasets import make_blobs

X, _ = make_blobs(n_samples=300, centers=4, cluster_std=0.7, random_state=0)

scores = {}
for k in range(2, 9):
    km = KMeans(n_clusters=k, random_state=0, n_init=10)
    labels = km.fit_predict(X)
    scores[k] = silhouette_score(X, labels)
    print(f'k={k}  silhouette={scores[k]:.3f}')

best_k = max(scores, key=scores.get)
print(f'Best k: {best_k}')

Silhouette Plots for Per-Point Analysis

A silhouette plot shows the silhouette coefficient of every individual point, sorted by cluster and width. Wide, uniform bars indicate all points are well-placed. Thin bars or points with negative scores reveal misassigned outliers. scikit-learn's silhouette_samples returns per-point scores that you can visualise this way.

from sklearn.metrics import silhouette_samples
import numpy as np

from sklearn.cluster import KMeans
from sklearn.datasets import make_blobs

X, _ = make_blobs(n_samples=100, centers=3, cluster_std=0.6, random_state=0)
km = KMeans(n_clusters=3, random_state=0, n_init=10)
labels = km.fit_predict(X)

samples = silhouette_samples(X, labels)
print('Per-cluster mean silhouettes:')
for c in range(3):
    print(f'  Cluster {c}: {samples[labels == c].mean():.3f}')

Combining Elbow and Silhouette

In practice, use both methods together. If the elbow suggests k=4 and the silhouette score is also highest at k=4, you have strong convergent evidence. When they disagree — e.g., elbow at k=3 but silhouette peaks at k=5 — examine the silhouette plot for each candidate k and apply domain knowledge to make the final call.

Gap Statistic: A Statistical Test for k

The gap statistic compares the observed inertia against the expected inertia under a null reference distribution (data sampled uniformly in the feature space). Choose the smallest k where gap(k) >= gap(k+1) - stddev. It is more statistically rigorous than the elbow method but computationally expensive because it requires generating many random reference datasets.

Practical Guidelines for k Selection

Start with domain knowledge — if you know there are 5 product categories, start with k=5. Use the elbow as a quick visual sanity check. Confirm with silhouette for objectivity. Evaluate downstream — for business use cases, test whether the segments are actionable and interpretable. The numerically optimal k is not always the most useful business segmentation.

Elbow and Silhouette Together: Full Example

Here is a compact pipeline that runs both diagnostics side by side, giving you a summary table to help pick k efficiently.

from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score
from sklearn.datasets import make_blobs
from sklearn.preprocessing import StandardScaler

X, _ = make_blobs(n_samples=400, centers=5, cluster_std=0.8, random_state=7)
X = StandardScaler().fit_transform(X)

print(f'{'k':>3}  {'Inertia':>10}  {'Silhouette':>10}')
for k in range(2, 10):
    km = KMeans(n_clusters=k, n_init=10, random_state=0)
    labels = km.fit_predict(X)
    sil = silhouette_score(X, labels)
    print(f'{k:>3}  {km.inertia_:>10.1f}  {sil:>10.3f}')

Quick Check

Test your understanding of k selection methods from this lesson.

Lesson Recap

In this lesson you learned: the elbow method plots inertia vs k and looks for the kink where improvement slows, silhouette score ranges from -1 to +1 and measures both cohesion and separation, and combining both methods with domain knowledge gives the most reliable k selection. Next up we explore DBSCAN — a density-based algorithm that discovers clusters of arbitrary shape and handles noise.

Często zadawane pytania

Czy lekcja „Wybór K: metoda łokcia i współczynnik sylwetki” jest bezpłatna?

Tak — pełny tekst „Wybór K: metoda łokcia i współczynnik sylwetki” jest dostępny za darmo tutaj w sieci. Aby ćwiczyć ją interaktywnie (wbudowany edytor kodu i tutor AI dostępny 24/7) i odblokować resztę kursu Machine Learning Academy, przejdź na CoddyKit PRO. Kurs Machine Learning Academy zawiera 4 lekcji w sumie.

Co nauczysz się w „Wybór K: metoda łokcia i współczynnik sylwetki”?

Uczą się Państwo rysować zależność inercji od k (metoda łokcia) oraz obliczać współczynniki sylwetki, aby wybrać liczbę klastrów tworzących zwarte i dobrze odseparowane grupy. Ćwiczysz Machine Learning Academy z praktycznym kodem, który uruchamiasz bezpośrednio w przeglądarce, a tutor AI dostępny 24/7 odpowiada na Twoje pytania podczas pracy nad lekcją.

Czy potrzebuję doświadczenia, aby zacząć Machine Learning Academy?

Nie wymagamy żadnego doświadczenia. Machine Learning Academy w CoddyKit jest strukturyzowany dla początkujących i zaawansowanych użytkowników, więc możesz zacząć tutaj lub od początku i uczyć się w swoim tempie. To lekcja 2 z 4.

Ile czasu zajmuje lekcja „Wybór K: metoda łokcia i współczynnik sylwetki”?

Większość lekcji CoddyKit trwa około 5–10 minut. Każda lekcja to mały, interaktywny krok, dzięki czemu robisz systematyczne postępy i zawsze wracasz dokładnie do tego samego miejsca — na webie i w aplikacji.

Czy mogę pisać i uruchamiać kod w tej lekcji Machine Learning Academy?

Tak. Każda lekcja Machine Learning Academy zawiera wbudowany edytor kodu, więc piszesz i uruchamiasz prawdziwy kod bezpośrednio w przeglądarce i od razu otrzymujesz sprzężenie zwrotne od AI — bez konfiguracji na komputerze.

Wszystkie lekcje w tym kursie

  1. K-Means: centroidy, przypisywanie i aktualizacja
  2. Wybór K: metoda łokcia i współczynnik sylwetki
  3. DBSCAN: punkty rdzeniowe, brzegowe i szum
  4. Klasteryzacja do segmentacji klientów: przykład kompleksowy
← Powrót do Machine Learning Academy