Machine Learning Academy · Leçon

Choisir K : méthode du coude et score de silhouette

Vous représenterez l’inertie en fonction de k (méthode du coude) et calculerez les coefficients de silhouette pour choisir le nombre de groupes qui produit des ensembles compacts et bien séparés.

Leçon 2 sur 413 étapes

Choisir K : méthode du coude et score de silhouette est une leçon Machine Learning Academy gratuite sur CoddyKit. Ceci est la leçon 2 sur 4. Tu peux lire la leçon complète ci-dessous gratuitement — puis la pratiquer en direct dans le navigateur avec un éditeur de code intégré et un tuteur IA 24/7. Elle fait partie du parcours d'apprentissage Machine Learning Academy, et ta progression se synchronise sur le web et l'application CoddyKit. Le cours Machine Learning Academy comprend 4 leçons au total.

Certaines parties de cette leçon n'ont pas encore été traduites et s'affichent en anglais.

Why Choosing k Matters

K-Means requires you to specify k — the number of clusters — before training. Too few clusters and you lump distinct groups together; too many and you split natural groups artificially. There is no universally correct k, but two diagnostic tools — the elbow method and the silhouette score — give principled guidance.

Inertia Decreases as k Grows

As you increase k, inertia always decreases because points are assigned to closer centroids. At k=n (one cluster per point), inertia is zero. This means you cannot simply minimise inertia — you need to find where additional clusters stop providing meaningful reductions. That point of diminishing returns is the elbow.

from sklearn.cluster import KMeans
import numpy as np

X = np.random.randn(200, 2)
inertias = []

for k in range(1, 11):
    km = KMeans(n_clusters=k, random_state=42, n_init=10)
    km.fit(X)
    inertias.append(km.inertia_)

print('Inertia per k:')
for k, inr in enumerate(inertias, start=1):
    print(f'  k={k}: {inr:.1f}')

The Elbow Method Explained

Plot inertia on the y-axis against k on the x-axis. The curve typically drops steeply for the first few k values then flattens. The elbow — the kink where the rate of decrease sharply slows — is your estimate of the true cluster count. If the true k is 3, the drop from k=1 to k=3 is large, but from k=3 to k=4 is much smaller.

import matplotlib.pyplot as plt
from sklearn.cluster import KMeans
from sklearn.datasets import make_blobs

X, _ = make_blobs(n_samples=300, centers=4, cluster_std=0.7, random_state=0)

inertias = []
for k in range(1, 11):
    km = KMeans(n_clusters=k, random_state=0, n_init=10)
    km.fit(X)
    inertias.append(km.inertia_)

plt.plot(range(1, 11), inertias, marker='o')
plt.xlabel('Number of clusters k')
plt.ylabel('Inertia')
plt.title('Elbow Method')
plt.axvline(x=4, color='red', linestyle='--', label='True k=4')
plt.legend()
plt.show()

Limitations of the Elbow Method

The elbow method works well when clusters are clearly separated, but real-world data often produces a smooth curve with no obvious kink. In such cases the elbow is ambiguous and different people may pick different k. That is where the silhouette score provides a more objective, mathematically grounded alternative.

Silhouette Score: The Formula

For each point i, compute two values: a(i) = mean distance to other points in the same cluster (cohesion), and b(i) = mean distance to the nearest different cluster (separation). The silhouette for point i is s(i) = (b(i) - a(i)) / max(a(i), b(i)). Values range from −1 (wrong cluster) through 0 (on border) to +1 (tight, well-separated cluster).

Computing Silhouette Score in sklearn

sklearn.metrics.silhouette_score returns the mean silhouette over all points. A score above 0.5 typically indicates reasonable clustering; above 0.7 is strong. Because you cannot compute the silhouette for k=1 (no second cluster), sweep k from 2 to some maximum and pick the k with the highest mean score.

from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score
from sklearn.datasets import make_blobs

X, _ = make_blobs(n_samples=300, centers=4, cluster_std=0.7, random_state=0)

scores = {}
for k in range(2, 9):
    km = KMeans(n_clusters=k, random_state=0, n_init=10)
    labels = km.fit_predict(X)
    scores[k] = silhouette_score(X, labels)
    print(f'k={k}  silhouette={scores[k]:.3f}')

best_k = max(scores, key=scores.get)
print(f'Best k: {best_k}')

Silhouette Plots for Per-Point Analysis

A silhouette plot shows the silhouette coefficient of every individual point, sorted by cluster and width. Wide, uniform bars indicate all points are well-placed. Thin bars or points with negative scores reveal misassigned outliers. scikit-learn's silhouette_samples returns per-point scores that you can visualise this way.

from sklearn.metrics import silhouette_samples
import numpy as np

from sklearn.cluster import KMeans
from sklearn.datasets import make_blobs

X, _ = make_blobs(n_samples=100, centers=3, cluster_std=0.6, random_state=0)
km = KMeans(n_clusters=3, random_state=0, n_init=10)
labels = km.fit_predict(X)

samples = silhouette_samples(X, labels)
print('Per-cluster mean silhouettes:')
for c in range(3):
    print(f'  Cluster {c}: {samples[labels == c].mean():.3f}')

Combining Elbow and Silhouette

In practice, use both methods together. If the elbow suggests k=4 and the silhouette score is also highest at k=4, you have strong convergent evidence. When they disagree — e.g., elbow at k=3 but silhouette peaks at k=5 — examine the silhouette plot for each candidate k and apply domain knowledge to make the final call.

Gap Statistic: A Statistical Test for k

The gap statistic compares the observed inertia against the expected inertia under a null reference distribution (data sampled uniformly in the feature space). Choose the smallest k where gap(k) >= gap(k+1) - stddev. It is more statistically rigorous than the elbow method but computationally expensive because it requires generating many random reference datasets.

Practical Guidelines for k Selection

Start with domain knowledge — if you know there are 5 product categories, start with k=5. Use the elbow as a quick visual sanity check. Confirm with silhouette for objectivity. Evaluate downstream — for business use cases, test whether the segments are actionable and interpretable. The numerically optimal k is not always the most useful business segmentation.

Elbow and Silhouette Together: Full Example

Here is a compact pipeline that runs both diagnostics side by side, giving you a summary table to help pick k efficiently.

from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score
from sklearn.datasets import make_blobs
from sklearn.preprocessing import StandardScaler

X, _ = make_blobs(n_samples=400, centers=5, cluster_std=0.8, random_state=7)
X = StandardScaler().fit_transform(X)

print(f'{'k':>3}  {'Inertia':>10}  {'Silhouette':>10}')
for k in range(2, 10):
    km = KMeans(n_clusters=k, n_init=10, random_state=0)
    labels = km.fit_predict(X)
    sil = silhouette_score(X, labels)
    print(f'{k:>3}  {km.inertia_:>10.1f}  {sil:>10.3f}')

Quick Check

Test your understanding of k selection methods from this lesson.

Lesson Recap

In this lesson you learned: the elbow method plots inertia vs k and looks for the kink where improvement slows, silhouette score ranges from -1 to +1 and measures both cohesion and separation, and combining both methods with domain knowledge gives the most reliable k selection. Next up we explore DBSCAN — a density-based algorithm that discovers clusters of arbitrary shape and handles noise.

Gratuit pour commencer

Apprends Python avec un tuteur IA — gratuit

Écris et exécute du vrai code dans ton navigateur, obtiens de l'aide instantanée d'un tuteur IA disponible 24h/24, et reprends là où tu t'es arrêté sur le web ou dans l'app.

Cours
30
Leçons
120

Questions Fréquemment Posées

La leçon « Choisir K : méthode du coude et score de silhouette » est-elle gratuite ?

Oui — le texte complet de « Choisir K : méthode du coude et score de silhouette » est gratuit à lire ici sur le web. Pour la pratiquer de manière interactive (un éditeur de code intégré et un tuteur IA 24/7) et déverrouiller le reste du cours Machine Learning Academy, passe à CoddyKit PRO. Le cours Machine Learning Academy comprend 4 leçons au total.

Qu'est-ce que j'apprendrai dans « Choisir K : méthode du coude et score de silhouette » ?

Vous représenterez l’inertie en fonction de k (méthode du coude) et calculerez les coefficients de silhouette pour choisir le nombre de groupes qui produit des ensembles compacts et bien séparés. Tu pratiques Machine Learning Academy avec du code pratique que tu exécutes directement dans le navigateur, et un tuteur IA 24/7 répond à tes questions au fur et à mesure que tu avances dans la leçon.

Dois-je avoir de l'expérience pour commencer Machine Learning Academy ?

Aucune expérience préalable n'est requise. Machine Learning Academy sur CoddyKit est structuré pour les débutants jusqu'aux apprenants avancés, donc tu peux commencer ici ou depuis le début et avancer à ton rythme. Ceci est la leçon 2 sur 4.

Combien de temps prend la leçon « Choisir K : méthode du coude et score de silhouette » ?

La plupart des leçons CoddyKit prennent environ 5–10 minutes. Chacune est courte et interactive, tu progresses régulièrement et tu repiques exactement où tu t'es arrêté sur le web et l'app.

Peux-tu écrire et exécuter du code dans cette leçon Machine Learning Academy ?

Oui. Chaque leçon Machine Learning Academy inclut un éditeur de code intégré, tu écris et exécutes du vrai code directement dans ton navigateur et tu reçois des retours IA instantanés — aucune configuration locale requise.

Toutes les leçons de ce cours

  1. K-Means : centroïdes, affectation et étapes de mise à jour
  2. Choisir K : méthode du coude et score de silhouette
  3. DBSCAN : points centraux, points de bord et bruit
  4. Regroupement pour segmenter les clients : exemple de bout en bout
← Retour à Machine Learning Academy