DBSCAN: 핵심 점, 경계 점 및 잡음
학습자는 eps와 min_samples를 설정하고, 초승달 모양 데이터셋에서 핵심·경계·잡음 점을 식별하며, K-Means가 놓치는 비볼록 클러스터를 DBSCAN이 찾아내는 과정을 확인합니다.
DBSCAN: 핵심 점, 경계 점 및 잡음은(는) CoddyKit의 무료 Machine Learning Academy 강의입니다. 이것은 4개 중 3번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 Machine Learning Academy 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. Machine Learning Academy 강의에는 총 4개의 강의가 포함되어 있습니다.
이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.
Why K-Means Fails on Arbitrary Shapes
K-Means assumes clusters are convex and roughly spherical. It fails on crescent, ring, or elongated shapes because it partitions by distance to centroids. DBSCAN (Density-Based Spatial Clustering of Applications with Noise) overcomes this by defining clusters as dense regions separated by low-density areas, discovering clusters of any shape.
Two Key Hyperparameters: eps and min_samples
DBSCAN is controlled by two parameters: eps (epsilon) defines the radius of a neighbourhood around a point, and min_samples sets the minimum number of points (including the point itself) required within that radius to be considered a dense region. Together they determine which points are cores, borders, or noise.
Core Points: The Anchors of Dense Regions
A point is a core point if at least min_samples points (including itself) lie within distance eps. Core points are the seeds from which clusters grow. Every point within the core's neighbourhood is directly reachable from it — the foundation for expanding the cluster.
from sklearn.neighbors import BallTree
import numpy as np
X = np.array([[0, 0], [0.3, 0], [0.6, 0],
[5, 5], [10, 10]])
eps = 1.0
min_samples = 3
tree = BallTree(X)
counts = tree.query_radius(X, r=eps, count_only=True)
core_mask = counts >= min_samples
print('Core points:', np.where(core_mask)[0]) # indices 0, 1, 2Border Points and Density-Reachability
A border point has fewer than min_samples neighbours within eps but lies within the eps-neighbourhood of a core point. It belongs to the cluster of its core point but does not expand the cluster further. A point is density-connected to another if there is a chain of directly-reachable steps linking them through core points.
Noise Points: Outlier Detection for Free
Points that are neither core nor border — isolated points with too few neighbours — are labelled noise (label = -1 in scikit-learn). This makes DBSCAN a natural outlier detector: anomalies that do not belong to any dense cluster are automatically flagged as noise without any extra configuration.
Running DBSCAN in scikit-learn
Use sklearn.cluster.DBSCAN. After fitting, db.labels_ contains integer cluster IDs starting at 0, with -1 for noise. db.core_sample_indices_ lists which samples are core points. The number of clusters is determined automatically — no k needed upfront.
from sklearn.cluster import DBSCAN
from sklearn.datasets import make_moons
import numpy as np
X, _ = make_moons(n_samples=200, noise=0.05, random_state=0)
db = DBSCAN(eps=0.3, min_samples=5)
db.fit(X)
n_clusters = len(set(db.labels_)) - (1 if -1 in db.labels_ else 0)
n_noise = (db.labels_ == -1).sum()
print('Clusters found:', n_clusters)
print('Noise points:', n_noise)
print('Labels (first 10):', db.labels_[:10])DBSCAN on Non-Convex Shapes
DBSCAN excels on datasets like two interlocking moons or concentric rings — shapes where K-Means completely fails. Because DBSCAN expands clusters along density chains, it naturally follows the curved manifold of the data. This is a fundamental algorithmic advantage for geospatial data, biological cell clusters, and anomaly-embedded datasets.
import matplotlib.pyplot as plt
from sklearn.cluster import DBSCAN, KMeans
from sklearn.datasets import make_moons
X, _ = make_moons(n_samples=300, noise=0.05, random_state=0)
db_labels = DBSCAN(eps=0.25, min_samples=5).fit_predict(X)
km_labels = KMeans(n_clusters=2, random_state=0, n_init=10).fit_predict(X)
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(10, 4))
ax1.scatter(X[:, 0], X[:, 1], c=db_labels, cmap='tab10')
ax1.set_title('DBSCAN')
ax2.scatter(X[:, 0], X[:, 1], c=km_labels, cmap='tab10')
ax2.set_title('K-Means')
plt.show()Choosing eps: The K-Distance Plot
A practical way to choose eps is to plot the k-distance graph: compute the distance of each point to its kth nearest neighbour (where k = min_samples), sort these distances, and look for the knee. The distance at the knee is a good eps candidate. Points above the knee are in sparse regions (noise); below the knee are in dense regions.
from sklearn.neighbors import NearestNeighbors
import numpy as np
import matplotlib.pyplot as plt
from sklearn.datasets import make_moons
X, _ = make_moons(n_samples=200, noise=0.05, random_state=0)
min_samples = 5
nn = NearestNeighbors(n_neighbors=min_samples)
nn.fit(X)
distances, _ = nn.kneighbors(X)
k_distances = np.sort(distances[:, -1])[::-1]
plt.plot(k_distances)
plt.xlabel('Points sorted by distance')
plt.ylabel(f'{min_samples}-th nearest neighbour distance')
plt.title('K-Distance Plot for eps selection')
plt.show()Effect of eps and min_samples on Results
Increasing eps merges clusters (eventually everything becomes one cluster). Decreasing eps creates more clusters and more noise. Increasing min_samples requires denser cores, making it harder to form clusters and generating more noise points. Tuning both parameters together is needed — the k-distance plot guides eps while min_samples is typically set to the dimensionality of the data plus one as a starting point.
DBSCAN vs K-Means: When to Use Each
Use DBSCAN when: clusters have irregular shapes, you do not know k in advance, outlier detection is important, or data has varying density. Use K-Means when: clusters are roughly spherical, the dataset is very large (DBSCAN scales as O(n log n) with a spatial index), or you need a specific number of clusters for business reasons like market segmentation into exactly 5 regions.
DBSCAN for Geospatial Clustering
DBSCAN is particularly popular for geospatial clustering (finding hotspots in GPS data) because it naturally identifies dense urban areas while marking sparse rural points as noise. Use metric='haversine' and convert coordinates to radians to cluster by great-circle distance on Earth's surface. The result is geographically meaningful clusters without needing to specify their number.
import numpy as np
from sklearn.cluster import DBSCAN
# Sample GPS coords: (lat, lon) in radians
coords = np.radians([
[40.7128, -74.0060], # NYC
[40.6892, -74.0445], # nearby
[40.7282, -73.7949], # Queens
[51.5074, -0.1278], # London
])
# eps in radians: 1km / earth radius
eps_rad = 1.0 / 6371.0
db = DBSCAN(eps=eps_rad, min_samples=2, metric='haversine')
db.fit(coords)
print('Cluster labels:', db.labels_)Quick Check
Test your understanding of DBSCAN concepts from this lesson.
Lesson Recap
In this lesson you learned: DBSCAN classifies points as core, border, or noise based on the eps radius and min_samples threshold, it discovers clusters of arbitrary shape by chaining density-reachable core points, and noise points (label -1) are automatic outliers — a feature K-Means does not provide. Next up we apply clustering end-to-end in a customer segmentation project.
AI 튜터와 함께 Python을(를) 배우세요 — 무료
브라우저에서 실제 코드를 작성하고 실행하며, 24/7 AI 튜터로부터 즉각적인 도움을 받고, 웹이나 앱에서 중단한 부분부터 계속 학습하세요.
- 코스
- 30
- 레슨
- 120
자주 묻는 질문
“DBSCAN: 핵심 점, 경계 점 및 잡음” 강의는 무료인가요?
네 — “DBSCAN: 핵심 점, 경계 점 및 잡음” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 Machine Learning Academy 강의 전체를 잠금 해제할 수 있습니다. Machine Learning Academy 강의에는 총 4개의 강의가 포함되어 있습니다.
“DBSCAN: 핵심 점, 경계 점 및 잡음”에서 뭘 배우나요?
학습자는 eps와 min_samples를 설정하고, 초승달 모양 데이터셋에서 핵심·경계·잡음 점을 식별하며, K-Means가 놓치는 비볼록 클러스터를 DBSCAN이 찾아내는 과정을 확인합니다. 브라우저에서 직접 실행하는 실습 코드로 Machine Learning Academy을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.
Machine Learning Academy을(를) 시작하는 데 경험이 필요한가요?
사전 경험은 필요하지 않습니다. CoddyKit의 Machine Learning Academy은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 3번째 강의입니다.
“DBSCAN: 핵심 점, 경계 점 및 잡음” 강의는 얼마나 걸리나요?
대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.
이 Machine Learning Academy 강의에서 코드를 작성하고 실행할 수 있나요?
네. 모든 Machine Learning Academy 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.
이 강의의 모든 강의
- K-Means: 중심점, 할당 및 갱신 단계
- K 선택하기: 엘보 방법과 실루엣 점수
- DBSCAN: 핵심 점, 경계 점 및 잡음
- 고객 세분화를 위한 클러스터링: 처음부터 끝까지 살펴보기