t-SNE: 시각화를 위한 이웃 구조 보존
학습자는 MNIST 임베딩에 서로 다른 perplexity 설정으로 t-SNE를 적용하고, t-SNE의 거리는 후속 모델링에서 의미가 없다는 점을 이해합니다.
t-SNE: 시각화를 위한 이웃 구조 보존은(는) CoddyKit의 무료 Machine Learning Academy 강의입니다. 이것은 4개 중 3번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 Machine Learning Academy 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. Machine Learning Academy 강의에는 총 4개의 강의가 포함되어 있습니다.
이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.
Beyond PCA: Non-Linear Visualisation
PCA projects data linearly and preserves global variance, but can fail to show local cluster structure. t-SNE (t-distributed Stochastic Neighbour Embedding) is a non-linear dimensionality reduction technique designed specifically for 2D and 3D visualisation. It prioritises preserving local neighbourhoods: points that are close in high-dimensional space should also be close in the 2D plot.
The Core Idea: Similarity Distributions
t-SNE defines a probability distribution over pairs of points in high-dimensional space: nearby points have high similarity. It then defines a similar distribution in the low-dimensional embedding. The algorithm minimises the KL divergence between the two distributions using gradient descent, nudging points in 2D until the neighbourhood structure matches the high-D structure.
The Perplexity Parameter
Perplexity is t-SNE's most important hyperparameter. It loosely controls how many neighbours each point considers when building the high-D similarity distribution — typically between 5 and 50. Low perplexity focuses on very local structure (many small clusters); high perplexity captures more global structure (broader, more spread-out clusters). The same dataset can look very different under different perplexity values.
Running t-SNE in scikit-learn
Use sklearn.manifold.TSNE. Key parameters: n_components (almost always 2), perplexity, n_iter (default 1000), and random_state. t-SNE is computationally expensive — O(n² log n) — so reduce the dataset with PCA first for large inputs (e.g., PCA to 50 dimensions, then t-SNE to 2D).
from sklearn.manifold import TSNE
from sklearn.datasets import load_digits
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
X, y = load_digits(return_X_y=True)
X_scaled = StandardScaler().fit_transform(X)
# Pre-reduce with PCA for speed
X_pca = PCA(n_components=30).fit_transform(X_scaled)
# t-SNE to 2D
tsne = TSNE(n_components=2, perplexity=30, n_iter=1000, random_state=42)
X_tsne = tsne.fit_transform(X_pca)
print('t-SNE shape:', X_tsne.shape)Visualising t-SNE Embeddings
Plot the 2D t-SNE coordinates with class colour to reveal cluster structure. On MNIST digits with appropriate perplexity, you typically see well-separated digit clusters, with similar-looking digits (e.g., 3 and 8) placed close together. This confirms that t-SNE captures semantically meaningful groupings.
import matplotlib.pyplot as plt
fig, ax = plt.subplots(figsize=(8, 6))
scatter = ax.scatter(X_tsne[:, 0], X_tsne[:, 1], c=y, cmap='tab10', s=8, alpha=0.7)
fig.colorbar(scatter, ax=ax, label='Digit')
ax.set_title('MNIST digits — t-SNE (perplexity=30)')
ax.set_xlabel('t-SNE 1')
ax.set_ylabel('t-SNE 2')
plt.tight_layout()
plt.show()Effect of Perplexity on the Embedding
It is essential to try multiple perplexity values and compare the plots. A perplexity that is too low creates many small disconnected blobs even within the same true cluster. A perplexity that is too high smears clusters together. A good practice is to test perplexity in [5, 15, 30, 50] and choose the embedding where known cluster structure appears most clearly.
import matplotlib.pyplot as plt
from sklearn.manifold import TSNE
perplexities = [5, 15, 30, 50]
fig, axes = plt.subplots(1, 4, figsize=(16, 4))
for ax, perp in zip(axes, perplexities):
tsne = TSNE(n_components=2, perplexity=perp, n_iter=800, random_state=0)
X_emb = tsne.fit_transform(X_pca[:300]) # subset for speed
ax.scatter(X_emb[:, 0], X_emb[:, 1], c=y[:300], cmap='tab10', s=10)
ax.set_title(f'Perplexity={perp}')
ax.axis('off')
plt.tight_layout()
plt.show()t-SNE Distances Are Not Meaningful
A critical warning: distances between clusters in t-SNE are not interpretable. A cluster appearing far from another does not mean they are globally distant; the algorithm optimises local neighbourhood preservation, not global distances. You cannot compare cluster sizes or inter-cluster distances across different runs or perplexity settings. Use t-SNE for exploration only, not for quantitative analysis.
t-SNE Is Stochastic and Non-Deterministic
Every t-SNE run with a different random_state produces a different layout — the embedding can rotate, reflect, or rearrange clusters. Always set random_state for reproducibility. Also, t-SNE does not have a transform method for out-of-sample points: you must refit on the entire dataset each time, which makes it unsuitable as a preprocessing step for a production model.
UMAP: A Modern Alternative to t-SNE
UMAP (Uniform Manifold Approximation and Projection) is a newer technique that is faster than t-SNE, preserves both local and more global structure, and supports transform for new points. It is not in scikit-learn but is installed via pip install umap-learn. For large datasets or production pipelines, UMAP is generally preferred over t-SNE.
# pip install umap-learn
import umap
reducer = umap.UMAP(n_components=2, n_neighbors=15, min_dist=0.1, random_state=42)
X_umap = reducer.fit_transform(X_pca)
import matplotlib.pyplot as plt
plt.scatter(X_umap[:, 0], X_umap[:, 1], c=y, cmap='tab10', s=8)
plt.title('MNIST — UMAP embedding')
plt.colorbar(label='Digit')
plt.show()When to Use t-SNE vs PCA
Use PCA for: preprocessing before modelling, compression, anomaly detection via reconstruction error, or when you need a deterministic, reversible transform. Use t-SNE for: exploring cluster structure in high-dimensional data, generating visualisations for presentations, or confirming that a dataset has meaningful groupings before applying a clustering or classification algorithm.
Complete t-SNE Visualisation Pipeline
Here is the recommended pipeline for t-SNE on any high-dimensional dataset: scale, reduce with PCA to ~50 dimensions, then apply t-SNE to 2D, and plot with class labels.
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
from sklearn.datasets import load_digits
import matplotlib.pyplot as plt
X, y = load_digits(return_X_y=True)
# Step 1: scale
X_s = StandardScaler().fit_transform(X)
# Step 2: PCA pre-reduction
X_pca = PCA(n_components=30, random_state=0).fit_transform(X_s)
# Step 3: t-SNE
X_tsne = TSNE(n_components=2, perplexity=30, random_state=0).fit_transform(X_pca)
# Step 4: plot
plt.scatter(X_tsne[:, 0], X_tsne[:, 1], c=y, cmap='tab10', s=10)
plt.title('Digits t-SNE')
plt.colorbar(label='Digit')
plt.show()Quick Check
Test your understanding of t-SNE from this lesson.
Lesson Recap
In this lesson you learned: t-SNE preserves local neighbourhoods by minimising KL divergence between high-D and low-D similarity distributions, perplexity controls the effective number of neighbours and should be tuned between 5 and 50, and t-SNE distances between clusters are not quantitatively meaningful — use it for exploration only. Next up we embed PCA inside a scikit-learn Pipeline as a preprocessing step for classifiers.
AI 튜터와 함께 Python을(를) 배우세요 — 무료
브라우저에서 실제 코드를 작성하고 실행하며, 24/7 AI 튜터로부터 즉각적인 도움을 받고, 웹이나 앱에서 중단한 부분부터 계속 학습하세요.
- 코스
- 30
- 레슨
- 120
자주 묻는 질문
“t-SNE: 시각화를 위한 이웃 구조 보존” 강의는 무료인가요?
네 — “t-SNE: 시각화를 위한 이웃 구조 보존” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 Machine Learning Academy 강의 전체를 잠금 해제할 수 있습니다. Machine Learning Academy 강의에는 총 4개의 강의가 포함되어 있습니다.
“t-SNE: 시각화를 위한 이웃 구조 보존”에서 뭘 배우나요?
학습자는 MNIST 임베딩에 서로 다른 perplexity 설정으로 t-SNE를 적용하고, t-SNE의 거리는 후속 모델링에서 의미가 없다는 점을 이해합니다. 브라우저에서 직접 실행하는 실습 코드로 Machine Learning Academy을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.
Machine Learning Academy을(를) 시작하는 데 경험이 필요한가요?
사전 경험은 필요하지 않습니다. CoddyKit의 Machine Learning Academy은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 3번째 강의입니다.
“t-SNE: 시각화를 위한 이웃 구조 보존” 강의는 얼마나 걸리나요?
대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.
이 Machine Learning Academy 강의에서 코드를 작성하고 실행할 수 있나요?
네. 모든 Machine Learning Academy 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.
이 강의의 모든 강의
- PCA: 분산, 고유벡터 및 주성분
- 데이터 투영 및 성분을 이용한 재구성
- t-SNE: 시각화를 위한 이웃 구조 보존
- 전처리로서의 PCA: 파이프라인에서 속도 향상과 잡음 감소