0Pricing
Machine Learning Academy · 강의

전처리로서의 PCA: 파이프라인에서 속도 향상과 잡음 감소

학습자는 분류기 앞의 sklearn Pipeline 안에 PCA를 포함하고, 차원 축소를 적용했을 때와 적용하지 않았을 때의 학습 시간과 테스트 정확도를 비교합니다.

전처리로서의 PCA: 파이프라인에서 속도 향상과 잡음 감소은(는) CoddyKit의 무료 Machine Learning Academy 강의입니다. 이것은 4개 중 4번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 Machine Learning Academy 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. Machine Learning Academy 강의에는 총 4개의 강의가 포함되어 있습니다.

이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.

PCA as a Preprocessing Step

Beyond visualisation, PCA serves as a practical preprocessing step that feeds compressed features into a downstream classifier or regressor. By discarding low-variance components that often encode noise, PCA can speed up training, reduce memory usage, and sometimes improve generalisation — especially when the original feature space is very high-dimensional.

Why PCA Can Reduce Noise

Random measurement noise typically spreads across many directions in feature space, but its variance is small in any single direction. PCA concentrates the meaningful signal in the top components and leaves noise in the low-variance tail. When you discard that tail, you effectively denoise the data. This is why PCA pre-processing sometimes helps algorithms like logistic regression that are sensitive to correlated or noisy features.

Embedding PCA in a sklearn Pipeline

The cleanest way to use PCA as preprocessing is inside a Pipeline. The pipeline ensures the scaler and PCA are fitted only on training data and then applied consistently to test data. This eliminates an entire class of data leakage bugs that occur when people forget to apply the same PCA transform to the test set.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split

X, y = load_digits(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

pipe = Pipeline([
    ('scaler', StandardScaler()),
    ('pca', PCA(n_components=0.95)),
    ('clf', LogisticRegression(max_iter=500))
])

pipe.fit(X_train, y_train)
print('Test accuracy:', pipe.score(X_test, y_test).round(4))

Comparing Training Time With and Without PCA

On high-dimensional datasets PCA can dramatically reduce training time because the classifier sees far fewer features. Let us benchmark logistic regression on the digits dataset (64 features) with and without PCA pre-reduction.

import time
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split

X, y = load_digits(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)

# Without PCA
t0 = time.time()
pipe_full = Pipeline([('sc', StandardScaler()), ('clf', LogisticRegression(max_iter=1000))])
pipe_full.fit(X_train, y_train)
t_full = time.time() - t0

# With PCA
t0 = time.time()
pipe_pca = Pipeline([('sc', StandardScaler()), ('pca', PCA(n_components=0.95)),
                     ('clf', LogisticRegression(max_iter=500))])
pipe_pca.fit(X_train, y_train)
t_pca = time.time() - t0

print(f'Without PCA: {t_full:.3f}s  acc={pipe_full.score(X_test, y_test):.4f}')
print(f'With PCA:    {t_pca:.3f}s  acc={pipe_pca.score(X_test, y_test):.4f}')

When PCA Helps and When It Does Not

PCA preprocessing helps most when: the number of features is large relative to the number of samples (high-dimensional, low-sample regime), features are correlated (redundant information), or the algorithm is slow with many features (e.g., SVM with RBF kernel). PCA tends NOT to help when: the dataset already has few, informative features, or you are using tree-based models (random forests, XGBoost) that handle redundant features natively and do not benefit from PCA's linear compression.

Tuning n_components in a Grid Search

Because PCA is a step inside a Pipeline, you can tune n_components alongside the classifier's hyperparameters using GridSearchCV. Use the double-underscore notation pca__n_components to refer to the PCA step's parameter. This lets the cross-validation loop find the optimal compression level simultaneously with the model's regularisation strength.

from sklearn.model_selection import GridSearchCV
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_digits

X, y = load_digits(return_X_y=True)

pipe = Pipeline([
    ('sc', StandardScaler()),
    ('pca', PCA()),
    ('clf', LogisticRegression(max_iter=500))
])

param_grid = {
    'pca__n_components': [10, 20, 30, 40],
    'clf__C': [0.1, 1.0, 10.0]
}

grid = GridSearchCV(pipe, param_grid, cv=5, n_jobs=-1)
grid.fit(X, y)
print('Best params:', grid.best_params_)
print('Best CV score:', grid.best_score_.round(4))

PCA Before SVM: A Classic Combination

SVMs with RBF kernels compute pairwise distances in the original feature space — expensive for high-dimensional data. Applying PCA first reduces dimensions while retaining signal, shrinking the distance computation. This combination was standard practice on image classification tasks before deep learning dominated: reduce image pixels with PCA, then classify with SVM.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
from sklearn.svm import SVC
from sklearn.datasets import load_digits
from sklearn.model_selection import cross_val_score
import numpy as np

X, y = load_digits(return_X_y=True)

pipe = Pipeline([
    ('sc', StandardScaler()),
    ('pca', PCA(n_components=30)),
    ('svm', SVC(kernel='rbf', C=10, gamma='scale'))
])

scores = cross_val_score(pipe, X, y, cv=5)
print(f'Accuracy: {np.mean(scores):.4f} +/- {np.std(scores):.4f}')

PCA for Noise Reduction: Concrete Example

Let us add Gaussian noise to the digits dataset and compare classifier accuracy with and without PCA denoising. On noisy data, PCA often improves accuracy by discarding the noise-dominated low-variance components.

import numpy as np
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_digits
from sklearn.model_selection import cross_val_score

X, y = load_digits(return_X_y=True)
X_noisy = X + np.random.randn(*X.shape) * 5.0  # heavy noise

for nc in [None, 10, 20, 30, 40]:
    steps = [('sc', StandardScaler())]
    if nc:
        steps.append(('pca', PCA(n_components=nc)))
    steps.append(('clf', LogisticRegression(max_iter=500)))
    pipe = Pipeline(steps)
    acc = cross_val_score(pipe, X_noisy, y, cv=5).mean()
    label = f'PCA({nc})' if nc else 'No PCA'
    print(f'{label:10s}  acc={acc:.4f}')

Always Fit PCA on Training Data Only

A critical rule: never fit the PCA transform on the test set. Fitting on test data leaks test-set statistics into the preprocessing and gives overly optimistic performance estimates. Using a Pipeline enforces this rule automatically: when you call pipeline.fit(X_train, y_train), every step — including PCA — is fitted only on the training split.

Memory Savings from PCA

On datasets with millions of samples and thousands of features (e.g., text TF-IDF matrices, genomic data), PCA dramatically reduces memory footprint. A 1M×5000 matrix at float32 costs 20 GB; PCA to 100 components gives a 1M×100 matrix at 400 MB — a 50× reduction. For such cases, use IncrementalPCA from scikit-learn, which fits PCA in chunks and never needs to load the full matrix into RAM.

from sklearn.decomposition import IncrementalPCA
import numpy as np

# Simulate large dataset as batches
n_samples, n_features = 10000, 500
n_components = 50
batch_size = 500

ipca = IncrementalPCA(n_components=n_components)
for i in range(0, n_samples, batch_size):
    batch = np.random.randn(batch_size, n_features)
    ipca.partial_fit(batch)

print('Explained variance ratio sum:', ipca.explained_variance_ratio_.sum().round(4))

Combining PCA with ColumnTransformer

In mixed-type datasets, you can apply PCA only to numeric columns while encoding categoricals separately, then combine with a ColumnTransformer. This is an advanced but realistic pattern for production ML pipelines where image features, numerical measurements, and categorical flags all coexist in the same row.

Quick Check

Test your understanding of PCA as pipeline preprocessing from this lesson.

Lesson Recap

In this lesson you learned: PCA inside a Pipeline prevents data leakage by fitting the transform only on training data, PCA can speed up training and reduce noise especially for high-dimensional or correlated feature sets, and n_components can be tuned via GridSearchCV alongside other model hyperparameters using double-underscore notation. Next up we build our first complete scikit-learn Pipeline combining scaler and classifier.

자주 묻는 질문

“전처리로서의 PCA: 파이프라인에서 속도 향상과 잡음 감소” 강의는 무료인가요?

네 — “전처리로서의 PCA: 파이프라인에서 속도 향상과 잡음 감소” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 Machine Learning Academy 강의 전체를 잠금 해제할 수 있습니다. Machine Learning Academy 강의에는 총 4개의 강의가 포함되어 있습니다.

“전처리로서의 PCA: 파이프라인에서 속도 향상과 잡음 감소”에서 뭘 배우나요?

학습자는 분류기 앞의 sklearn Pipeline 안에 PCA를 포함하고, 차원 축소를 적용했을 때와 적용하지 않았을 때의 학습 시간과 테스트 정확도를 비교합니다. 브라우저에서 직접 실행하는 실습 코드로 Machine Learning Academy을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.

Machine Learning Academy을(를) 시작하는 데 경험이 필요한가요?

사전 경험은 필요하지 않습니다. CoddyKit의 Machine Learning Academy은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 4번째 강의입니다.

“전처리로서의 PCA: 파이프라인에서 속도 향상과 잡음 감소” 강의는 얼마나 걸리나요?

대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.

이 Machine Learning Academy 강의에서 코드를 작성하고 실행할 수 있나요?

네. 모든 Machine Learning Academy 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.

이 강의의 모든 강의

  1. PCA: 분산, 고유벡터 및 주성분
  2. 데이터 투영 및 성분을 이용한 재구성
  3. t-SNE: 시각화를 위한 이웃 구조 보존
  4. 전처리로서의 PCA: 파이프라인에서 속도 향상과 잡음 감소
← Machine Learning Academy(으)로 돌아가기