0Pricing
Data Science Academy · Lezione

k-Means e scelta di k

I metodi del gomito e della silhouette.

k-Means e scelta di k è una lezione Data Science Academy gratuita su CoddyKit. Questa è la lezione 2 di 4. Puoi leggere la lezione completa qui gratuitamente — poi esercitati direttamente nel browser con un editor di codice integrato e un tutor IA disponibile 24/7. Fa parte del percorso di apprendimento Data Science Academy, e i tuoi progressi si sincronizzano tra il web e l'app CoddyKit. Il corso Data Science Academy include 4 lezioni in totale.

Parti di questa lezione non sono ancora state tradotte e vengono mostrate in inglese.

The Go-To Clusterer

k-Means is the most popular clustering algorithm: it splits your data into k groups by grouping points around central anchors. 🎯

Centroids Are Anchors

Each cluster has a centroid, the mean position of its members. Points join the cluster whose centroid sits closest to them.

The Repeating Dance

k-Means alternates two steps: assign every point to its nearest centroid, then recompute each centroid from its new members.

It Settles Down

Those steps repeat until assignments stop changing. The algorithm has converged, and the clusters are final.

You Choose k First

k-Means cannot invent the number of groups. You must pick k up front, and the right value is rarely obvious.

Run It in One Line

scikit-learn makes fitting easy. Set n_clusters to your chosen k and call fit on your scaled data.

from sklearn.cluster import KMeans
model = KMeans(n_clusters=3).fit(X)

Inertia Measures Tightness

After fitting, inertia reports the total squared distance from points to their centroids. Lower means tighter clusters.

model.inertia_

The Elbow Method

Plot inertia for many k values. Where the curve bends sharply, the elbow, often marks a good number of clusters.

Diminishing Returns

Past the elbow, adding clusters barely lowers inertia. That flattening signals you are just splitting groups that already fit well.

The Silhouette Score

A second guide is the silhouette score: it rewards points that sit close to their own cluster and far from others.

from sklearn.metrics import silhouette_score

Start Smarter

Random starts can land badly, so scikit-learn defaults to k-means++, which spreads initial centroids out for steadier results.

Quick Check

Let us check how you would pick a sensible k.

Recap

k-Means groups points around centroids for a chosen k; the elbow and silhouette help you pick that k well. 🎯

Domande Frequenti

La lezione «k-Means e scelta di k» è gratuita?

Sì — il testo completo di «k-Means e scelta di k» è gratuito qui sul web. Per esercitarvi in modo interattivo (un editor di codice integrato e un tutor IA 24/7) e sbloccare il resto del corso Data Science Academy, passa a CoddyKit PRO. Il corso Data Science Academy include 4 lezioni in totale.

Cosa imparerò in «k-Means e scelta di k»?

I metodi del gomito e della silhouette. Eserciti Data Science Academy con codice pratico che esegui direttamente nel browser, e un tutor IA 24/7 risponde alle tue domande mentre lavori sulla lezione.

Ho bisogno di esperienza per iniziare Data Science Academy?

Non è richiesta alcuna esperienza precedente. Data Science Academy su CoddyKit è strutturato per principianti e studenti avanzati, quindi puoi iniziare da qui o dall'inizio e procedere al tuo ritmo. Questa è la lezione 2 di 4.

Quanto tempo richiede la lezione «k-Means e scelta di k»?

La maggior parte delle lezioni CoddyKit richiede circa 5–10 minuti. Ogni lezione è breve e interattiva, quindi fai progressi costanti e riprendi esattamente da dove hai lasciato su web e app.

Posso scrivere ed eseguire codice in questa lezione Data Science Academy?

Sì. Ogni lezione Data Science Academy include un editor di codice integrato, quindi scrivi ed esegui codice reale direttamente nel tuo browser e ricevi feedback istantaneo dall'IA — nessuna configurazione locale necessaria.

Tutte le lezioni di questo corso

  1. Apprendimento supervisionato e non supervisionato
  2. k-Means e scelta di k
  3. Clustering gerarchico e DBSCAN
  4. Profilare e nominare i cluster
← Torna a Data Science Academy