K-Means: 중심점, 할당 및 갱신 단계
학습자는 세 번의 K-Means 반복을 손으로 따라가며, 점을 가장 가까운 중심점에 할당하고 중심점을 다시 계산한 뒤 2차원 산점도에서 수렴 과정을 확인합니다.
K-Means: 중심점, 할당 및 갱신 단계은(는) CoddyKit의 무료 Machine Learning Academy 강의입니다. 이것은 4개 중 1번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 Machine Learning Academy 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. Machine Learning Academy 강의에는 총 4개의 강의가 포함되어 있습니다.
이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.
What Is K-Means Clustering?
K-Means is an unsupervised algorithm that partitions n data points into k non-overlapping clusters. Unlike supervised learning, there are no labels — the algorithm discovers structure purely from the feature values. K-Means is fast, scalable, and widely used for customer segmentation, image compression, and anomaly detection.
The Three-Step Algorithm
K-Means repeats three steps until convergence: 1) Initialise — randomly place k centroids in feature space. 2) Assignment — assign every point to the nearest centroid. 3) Update — move each centroid to the mean of its assigned points. The loop stops when assignments no longer change.
Computing Distance to Centroids
In each assignment step, the Euclidean distance from every point to every centroid is computed. A point is assigned to the centroid with the smallest distance. For a point x and centroid c, the squared distance is sum((x_i - c_i)^2). Using squared distance avoids the expensive square-root and gives the same ordering.
import numpy as np
def assign_clusters(X, centroids):
# X: (n, d), centroids: (k, d)
distances = np.linalg.norm(X[:, np.newaxis] - centroids, axis=2) # (n, k)
return np.argmin(distances, axis=1) # label for each point
X = np.array([[1, 2], [3, 4], [5, 6], [8, 8]])
centroids = np.array([[2, 2], [7, 7]])
labels = assign_clusters(X, centroids)
print(labels) # [0, 0, 0, 1]The Update Step: Recomputing Centroids
After assignment, each centroid is relocated to the arithmetic mean of all points currently in its cluster. If a cluster becomes empty (no points assigned), the centroid is usually re-initialised randomly or removed. This mean-shift minimises the total within-cluster sum of squares (WCSS) — also called inertia.
import numpy as np
def update_centroids(X, labels, k):
d = X.shape[1]
new_centroids = np.zeros((k, d))
for c in range(k):
points = X[labels == c]
if len(points) > 0:
new_centroids[c] = points.mean(axis=0)
return new_centroids
X = np.array([[1, 2], [3, 4], [5, 6], [8, 8]])
labels = np.array([0, 0, 0, 1])
print(update_centroids(X, labels, k=2))Tracing Convergence by Hand
Consider four 1D points: 1, 2, 8, 9 and k=2. Init: centroids = [1, 8]. Iter 1 assignment: 1→C0, 2→C0, 8→C1, 9→C1. Iter 1 update: C0=1.5, C1=8.5. Iter 2 assignment: unchanged. Converged in 2 iterations! In higher dimensions convergence may take more steps, but the logic is identical.
Inertia: Measuring Cluster Compactness
Inertia (WCSS) is the sum of squared distances between each point and its cluster centroid. Lower inertia means tighter, more compact clusters. K-Means minimises inertia at each update step, but the algorithm is not guaranteed to find the global minimum — it can get stuck in local optima depending on initialisation.
from sklearn.cluster import KMeans
import numpy as np
X = np.array([[1, 2], [1, 4], [1, 0],
[10, 2], [10, 4], [10, 0]])
km = KMeans(n_clusters=2, random_state=42)
km.fit(X)
print('Inertia:', km.inertia_)
print('Labels:', km.labels_)
print('Centroids:', km.cluster_centers_)K-Means++ Initialisation
Random centroid initialisation often leads to slow convergence or poor local optima. K-Means++ (the scikit-learn default via init='k-means++') seeds centroids more intelligently: the first centroid is chosen randomly, and each subsequent centroid is selected with probability proportional to its squared distance from the nearest already-chosen centroid. This spreads starting points and consistently finds better solutions.
from sklearn.cluster import KMeans
import numpy as np
X = np.random.randn(300, 2)
# Default: k-means++ initialisation
km = KMeans(n_clusters=3, init='k-means++', n_init=10, random_state=0)
km.fit(X)
print('Inertia with k-means++:', round(km.inertia_, 2))
# Compare with random init
km_rand = KMeans(n_clusters=3, init='random', n_init=10, random_state=0)
km_rand.fit(X)
print('Inertia with random init:', round(km_rand.inertia_, 2))Visualising Cluster Assignments
Plotting cluster assignments on a 2D scatter shows the Voronoi partition — the decision boundaries where each point's nearest centroid changes. Plotting centroids as large stars and colouring points by cluster label makes convergence intuitive. This visualisation also reveals when clusters overlap or have unequal sizes.
import matplotlib.pyplot as plt
from sklearn.cluster import KMeans
from sklearn.datasets import make_blobs
X, _ = make_blobs(n_samples=300, centers=3, cluster_std=0.6, random_state=0)
km = KMeans(n_clusters=3, random_state=0)
labels = km.fit_predict(X)
plt.scatter(X[:, 0], X[:, 1], c=labels, cmap='tab10', s=30)
plt.scatter(km.cluster_centers_[:, 0], km.cluster_centers_[:, 1],
c='black', s=200, marker='*', label='Centroids')
plt.legend()
plt.title('K-Means Clusters')
plt.show()Multiple Restarts and n_init
Because K-Means can converge to local optima, scikit-learn runs the algorithm n_init times with different random seeds and keeps the result with the lowest inertia. The default is n_init=10. For small datasets 10 is usually sufficient; for large or tricky datasets you may raise it to 20 or 50. Always check the final inertia against the best-run inertia to diagnose poor convergence.
from sklearn.cluster import KMeans
import numpy as np
X = np.random.randn(500, 5)
km = KMeans(n_clusters=4, n_init=20, random_state=0)
km.fit(X)
print('Best inertia over 20 runs:', round(km.inertia_, 2))
print('Number of iterations until convergence:', km.n_iter_)Limitations of K-Means
K-Means has several well-known weaknesses: 1) Assumes spherical clusters — it struggles with elongated or crescent shapes. 2) Sensitive to outliers — a distant outlier pulls the centroid away from the true cluster mean. 3) Requires k upfront — you must know or estimate the number of clusters before fitting. 4) Feature scale matters — always standardise features before running K-Means so large-scale variables do not dominate distances.
Running K-Means with scikit-learn
In practice, using sklearn.cluster.KMeans is the standard approach. Key parameters: n_clusters (k), init (default 'k-means++'), n_init, max_iter (default 300), and random_state. After fitting, km.labels_ contains cluster assignments, km.cluster_centers_ holds centroid positions, and km.inertia_ reports WCSS.
from sklearn.cluster import KMeans
from sklearn.preprocessing import StandardScaler
from sklearn.datasets import load_iris
X, _ = load_iris(return_X_y=True)
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
km = KMeans(n_clusters=3, random_state=42)
km.fit(X_scaled)
print('Cluster sizes:', {i: (km.labels_ == i).sum() for i in range(3)})
print('Inertia:', round(km.inertia_, 2))Quick Check
Test your understanding of K-Means clustering concepts from this lesson.
Lesson Recap
In this lesson you learned: K-Means iterates assignment and update steps until cluster memberships stabilise, inertia (WCSS) measures compactness and is minimised by each update, and K-Means++ initialisation and multiple restarts help avoid poor local optima. Next up we explore how to choose the right value of k using the elbow method and silhouette score.
자주 묻는 질문
“K-Means: 중심점, 할당 및 갱신 단계” 강의는 무료인가요?
네 — “K-Means: 중심점, 할당 및 갱신 단계” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 Machine Learning Academy 강의 전체를 잠금 해제할 수 있습니다. Machine Learning Academy 강의에는 총 4개의 강의가 포함되어 있습니다.
“K-Means: 중심점, 할당 및 갱신 단계”에서 뭘 배우나요?
학습자는 세 번의 K-Means 반복을 손으로 따라가며, 점을 가장 가까운 중심점에 할당하고 중심점을 다시 계산한 뒤 2차원 산점도에서 수렴 과정을 확인합니다. 브라우저에서 직접 실행하는 실습 코드로 Machine Learning Academy을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.
Machine Learning Academy을(를) 시작하는 데 경험이 필요한가요?
사전 경험은 필요하지 않습니다. CoddyKit의 Machine Learning Academy은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 1번째 강의입니다.
“K-Means: 중심점, 할당 및 갱신 단계” 강의는 얼마나 걸리나요?
대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.
이 Machine Learning Academy 강의에서 코드를 작성하고 실행할 수 있나요?
네. 모든 Machine Learning Academy 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.
이 강의의 모든 강의
- K-Means: 중심점, 할당 및 갱신 단계
- K 선택하기: 엘보 방법과 실루엣 점수
- DBSCAN: 핵심 점, 경계 점 및 잡음
- 고객 세분화를 위한 클러스터링: 처음부터 끝까지 살펴보기