Machine Learning Academy · 강의

PCA: 분산, 고유벡터 및 주성분

학습자는 고차원 데이터셋에 PCA를 적합하고, 설명된 분산 비율을 확인하며, 전체 분산의 95%를 보존하는 성분 수를 선택합니다.

레슨 1/413개 단계

PCA: 분산, 고유벡터 및 주성분은(는) CoddyKit의 무료 Machine Learning Academy 강의입니다. 이것은 4개 중 1번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 Machine Learning Academy 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. Machine Learning Academy 강의에는 총 4개의 강의가 포함되어 있습니다.

이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.

The Problem with High-Dimensional Data

As feature count grows, datasets become increasingly sparse — the curse of dimensionality. Many features are redundant or correlated, carrying overlapping information. Principal Component Analysis (PCA) solves this by finding a new, smaller set of axes (principal components) that capture the maximum variance in the data with the fewest dimensions.

Variance: What PCA Maximises

PCA seeks directions in feature space along which the data varies the most. A direction with high variance captures rich information; a direction with near-zero variance is essentially noise. The first principal component (PC1) is the direction of maximum variance, PC2 is orthogonal to PC1 with the next highest variance, and so on.

Covariance Matrix and Eigenvectors

PCA operates on the covariance matrix of the centred data. The eigenvectors of this matrix point in the directions of maximum variance, and the corresponding eigenvalues measure how much variance each direction captures. The eigenvectors are the principal components; sorting them by eigenvalue in descending order gives PC1, PC2, ... PCn.

import numpy as np

X = np.array([[2.5, 2.4], [0.5, 0.7], [2.2, 2.9],
              [1.9, 2.2], [3.1, 3.0], [2.3, 2.7]])

# Centre the data
X_centered = X - X.mean(axis=0)

# Compute covariance matrix
cov = np.cov(X_centered.T)
print('Covariance matrix:\n', cov)

# Eigenvectors and eigenvalues
eigenvalues, eigenvectors = np.linalg.eigh(cov)
idx = np.argsort(eigenvalues)[::-1]
print('Eigenvalues:', eigenvalues[idx])
print('PC1 direction:', eigenvectors[:, idx[0]])

Explained Variance Ratio

The explained variance ratio of each component is its eigenvalue divided by the sum of all eigenvalues. If PC1 explains 90% of variance and PC2 explains 8%, the first two components together retain 98% of all information. This ratio guides how many components to keep — a common threshold is 95%.

from sklearn.decomposition import PCA
from sklearn.datasets import load_digits

X, _ = load_digits(return_X_y=True)  # 64 features

pca = PCA()
pca.fit(X)

cumulative_variance = pca.explained_variance_ratio_.cumsum()
n_95 = (cumulative_variance < 0.95).sum() + 1

print(f'Components to retain 95% variance: {n_95}')
print(f'Explained by first 10 components: {cumulative_variance[9]:.3f}')

Choosing n_components

Set n_components as an integer (e.g., PCA(n_components=10)) to keep exactly 10 components, or as a float between 0 and 1 (e.g., PCA(n_components=0.95)) to automatically keep enough components to explain that fraction of variance. The latter is the cleanest approach for pipelines where you want variance-based truncation without knowing the count upfront.

from sklearn.decomposition import PCA
from sklearn.datasets import load_digits

X, _ = load_digits(return_X_y=True)

# Retain 95% of variance automatically
pca = PCA(n_components=0.95)
pca.fit(X)

print('Number of components chosen:', pca.n_components_)
print('Total variance retained:', pca.explained_variance_ratio_.sum().round(4))

The Scree Plot

A scree plot shows explained variance ratio (or eigenvalue) on the y-axis and component index on the x-axis. The plot typically shows a steep drop then a flat plateau. The elbow — where the drop becomes gradual — is another heuristic for the number of components to retain, similar to the elbow method in K-Means.

import matplotlib.pyplot as plt
from sklearn.decomposition import PCA
from sklearn.datasets import load_wine

X, _ = load_wine(return_X_y=True)
pca = PCA()
pca.fit(X)

plt.figure(figsize=(8, 4))
plt.subplot(1, 2, 1)
plt.bar(range(1, 14), pca.explained_variance_ratio_)
plt.xlabel('Component')
plt.ylabel('Explained variance ratio')
plt.title('Scree Plot')
plt.subplot(1, 2, 2)
plt.plot(pca.explained_variance_ratio_.cumsum(), marker='o')
plt.axhline(0.95, color='red', linestyle='--')
plt.xlabel('Number of components')
plt.ylabel('Cumulative variance')
plt.tight_layout()
plt.show()

Centering and Scaling Before PCA

PCA is sensitive to feature scale. A feature measured in thousands will dominate the covariance matrix. Always standardise with StandardScaler before PCA to give each feature unit variance. Centring (zero mean) is essential — PCA implicitly does this, but if you use a Pipeline, the scaler should come first so PCA operates on already-centred, equal-scale features.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
from sklearn.datasets import load_wine

X, _ = load_wine(return_X_y=True)

pipe = Pipeline([
    ('scaler', StandardScaler()),
    ('pca', PCA(n_components=0.95))
])
pipe.fit(X)

print('Original shape:', X.shape)
print('Reduced shape:', pipe.transform(X).shape)

What Do Principal Components Represent?

Each principal component is a linear combination of the original features — a weighted sum. Inspecting the component loadings (the coefficients) reveals which original features contribute most to each PC. However, components are often not directly interpretable because they mix features together. PCA is primarily a compression tool, not a feature selection tool.

from sklearn.decomposition import PCA
from sklearn.preprocessing import StandardScaler
from sklearn.datasets import load_wine
import pandas as pd

X, _ = load_wine(return_X_y=True)
X_scaled = StandardScaler().fit_transform(X)

pca = PCA(n_components=2)
pca.fit(X_scaled)

feature_names = load_wine().feature_names
loadings = pd.DataFrame(pca.components_.T, index=feature_names,
                        columns=['PC1', 'PC2'])
print(loadings.round(2))

SVD: The Efficient Implementation

In practice, scikit-learn computes PCA via Singular Value Decomposition (SVD) rather than explicit eigendecomposition of the covariance matrix, because SVD is numerically more stable and works directly on the data matrix without forming the covariance matrix. The result is mathematically identical. For very large datasets, PCA(svd_solver='randomized') uses an approximate randomised SVD for speed.

PCA Is Linear and Orthogonal

Important limitations: PCA finds only linear relationships between features. If the meaningful structure in your data lies on a curved manifold (e.g., a Swiss roll), PCA will not discover it effectively — kernel PCA or t-SNE are better alternatives. Also, PCA components are orthogonal by construction, which can be a mismatch if your underlying factors are correlated.

PCA on a Real Dataset: Quick End-to-End

Here is the full workflow: scale, PCA to 2D, and scatter-plot with class colour to check if the reduced space still separates classes visually. This is a standard exploratory step before training a classifier on the full feature set.

from sklearn.decomposition import PCA
from sklearn.preprocessing import StandardScaler
from sklearn.datasets import load_iris
import matplotlib.pyplot as plt

X, y = load_iris(return_X_y=True)
X_scaled = StandardScaler().fit_transform(X)

pca = PCA(n_components=2)
X_2d = pca.fit_transform(X_scaled)

plt.scatter(X_2d[:, 0], X_2d[:, 1], c=y, cmap='Set1', s=30)
plt.xlabel(f'PC1 ({pca.explained_variance_ratio_[0]:.1%} var)')
plt.ylabel(f'PC2 ({pca.explained_variance_ratio_[1]:.1%} var)')
plt.title('Iris in PCA space')
plt.colorbar(label='Class')
plt.show()

Quick Check

Test your understanding of PCA from this lesson.

Lesson Recap

In this lesson you learned: PCA finds directions of maximum variance via the covariance matrix eigenvectors, explained variance ratio guides how many components to keep (typically aim for 95%), and always standardise features before PCA so scale differences do not bias the components. Next up we project data into principal-component space and reconstruct it to quantify information loss.

무료로 시작

AI 튜터와 함께 Python을(를) 배우세요 — 무료

브라우저에서 실제 코드를 작성하고 실행하며, 24/7 AI 튜터로부터 즉각적인 도움을 받고, 웹이나 앱에서 중단한 부분부터 계속 학습하세요.

코스
30
레슨
120

자주 묻는 질문

“PCA: 분산, 고유벡터 및 주성분” 강의는 무료인가요?

네 — “PCA: 분산, 고유벡터 및 주성분” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 Machine Learning Academy 강의 전체를 잠금 해제할 수 있습니다. Machine Learning Academy 강의에는 총 4개의 강의가 포함되어 있습니다.

“PCA: 분산, 고유벡터 및 주성분”에서 뭘 배우나요?

학습자는 고차원 데이터셋에 PCA를 적합하고, 설명된 분산 비율을 확인하며, 전체 분산의 95%를 보존하는 성분 수를 선택합니다. 브라우저에서 직접 실행하는 실습 코드로 Machine Learning Academy을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.

Machine Learning Academy을(를) 시작하는 데 경험이 필요한가요?

사전 경험은 필요하지 않습니다. CoddyKit의 Machine Learning Academy은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 1번째 강의입니다.

“PCA: 분산, 고유벡터 및 주성분” 강의는 얼마나 걸리나요?

대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.

이 Machine Learning Academy 강의에서 코드를 작성하고 실행할 수 있나요?

네. 모든 Machine Learning Academy 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.

이 강의의 모든 강의

  1. PCA: 분산, 고유벡터 및 주성분
  2. 데이터 투영 및 성분을 이용한 재구성
  3. t-SNE: 시각화를 위한 이웃 구조 보존
  4. 전처리로서의 PCA: 파이프라인에서 속도 향상과 잡음 감소
← Machine Learning Academy(으)로 돌아가기