Machine Learning Academy · レッスン

最初のPipelineを構築する:スケーラーと分類器

StandardScalerとLogisticRegressionをPipelineに連結し、fitとpredictを呼び出して、スケーラーが学習データだけでfitされていることを確認します。

レッスン 1/413 ステップ

「最初のPipelineを構築する:スケーラーと分類器」はCoddyKit上の無料Machine Learning Academyレッスンです。 これはレッスン1/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはMachine Learning Academy学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Machine Learning Academyコースには全4レッスンが含まれています。

このレッスンの一部はまだ翻訳されておらず、英語で表示されています。

What Is a scikit-learn Pipeline?

A Pipeline chains multiple processing steps into a single estimator object. Each step except the last must implement fit and transform; the final step only needs fit and predict. Calling pipeline.fit(X, y) runs all steps in sequence, and pipeline.predict(X) passes data through all transforms before classifying. This eliminates bookkeeping errors and prevents data leakage.

Why Pipelines Prevent Leakage

If you fit a StandardScaler on the full dataset and then split into train/test, the scaler has seen test-set statistics — this is data leakage. A Pipeline solves this automatically: when you pass a Pipeline to cross_val_score or GridSearchCV, the entire pipeline (including the scaler) is re-fitted from scratch on each training fold, so the test fold never influences the scaler parameters.

Constructing Your First Pipeline

Create a Pipeline by passing a list of (name, estimator) tuples. The names are arbitrary strings you choose — they are used to reference steps later (e.g., for grid-search hyperparameters). The most common first pipeline is a scaler followed by a classifier.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

pipe = Pipeline([
    ('scaler', StandardScaler()),
    ('clf', LogisticRegression(C=1.0, max_iter=200))
])

print('Steps:', [name for name, _ in pipe.steps])
print('Named steps:', list(pipe.named_steps.keys()))

Fitting and Predicting

After construction, use the Pipeline exactly like any sklearn estimator: fit on training data, predict on test data, score for accuracy. Internally, fit calls fit_transform on all intermediate steps and fit on the final estimator.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split

X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

pipe = Pipeline([
    ('scaler', StandardScaler()),
    ('clf', LogisticRegression(max_iter=300))
])

pipe.fit(X_train, y_train)
print('Test accuracy:', pipe.score(X_test, y_test).round(4))

y_pred = pipe.predict(X_test)
print('Predictions[:5]:', y_pred[:5])

Confirming Scaler Fitted Only on Train Data

After fitting the Pipeline on training data, you can inspect each step's fitted parameters. The scaler inside the pipeline will have mean_ and scale_ attributes computed from the training set only — not the full dataset. This confirms the pipeline is doing the right thing.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
import numpy as np

X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0)

pipe = Pipeline([('sc', StandardScaler()), ('lr', LogisticRegression(max_iter=300))])
pipe.fit(X_train, y_train)

# Scaler mean from pipeline vs computed from X_train
print('Pipeline scaler mean[0]:', pipe.named_steps['sc'].mean_[0].round(4))
print('Direct train mean[0]:   ', X_train[:, 0].mean().round(4))

Accessing Intermediate Outputs

To get the transformed output of a specific step, use pipeline[:-1].transform(X) or access individual steps via pipeline.named_steps['step_name']. You can also call pipeline[:'step_name'] with Python slice notation to get a sub-pipeline up to and including that step. This is useful for debugging or inspecting what the data looks like after scaling.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_iris

X, y = load_iris(return_X_y=True)

pipe = Pipeline([('sc', StandardScaler()), ('lr', LogisticRegression())])
pipe.fit(X, y)

# Get scaled output (all steps except the last)
X_scaled = pipe[:-1].transform(X)
print('After scaling — mean per feature:', X_scaled.mean(axis=0).round(4))
print('After scaling — std per feature:', X_scaled.std(axis=0).round(4))

make_pipeline: A Shortcut

make_pipeline creates a Pipeline without requiring you to specify step names manually — it generates names from the class names in lowercase. This is convenient for quick experiments but less readable in production code where explicit names help identify steps in grid-search parameter strings.

from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import MinMaxScaler
from sklearn.neighbors import KNeighborsClassifier
from sklearn.datasets import load_iris
from sklearn.model_selection import cross_val_score
import numpy as np

X, y = load_iris(return_X_y=True)

# Auto-names: 'minmaxscaler' and 'kneighborsclassifier'
pipe = make_pipeline(MinMaxScaler(), KNeighborsClassifier(n_neighbors=5))
print('Step names:', list(pipe.named_steps.keys()))
print('CV accuracy:', cross_val_score(pipe, X, y, cv=5).mean().round(4))

Pipeline with Probability Predictions

If the final estimator supports predict_proba, the Pipeline exposes it too. This allows you to use a Pipeline with any function that expects probability outputs — like ROC-AUC scoring, calibration curves, or threshold tuning — without needing to manually apply preprocessing before calling the model.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split

X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0)

pipe = Pipeline([('sc', StandardScaler()), ('lr', LogisticRegression(max_iter=300))])
pipe.fit(X_train, y_train)

proba = pipe.predict_proba(X_test)[:, 1]
print('ROC-AUC:', roc_auc_score(y_test, proba).round(4))

Setting Parameters After Construction

You can update any step parameter after building the Pipeline using set_params(step__param=value) with the double-underscore separator. This is the same syntax used in GridSearchCV. It is useful when you want to experiment with different settings without rebuilding the entire pipeline from scratch.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

pipe = Pipeline([('sc', StandardScaler()), ('lr', LogisticRegression())])

# Change C and max_iter after construction
pipe.set_params(lr__C=10.0, lr__max_iter=500)
print('C after set_params:', pipe.named_steps['lr'].C)

Pickling and Sharing Pipelines

A fitted Pipeline is a single Python object that can be serialised with pickle or joblib. Sharing the pipeline as one file ensures that the exact same preprocessing steps — with the exact same scaler parameters — are applied at prediction time, eliminating consistency bugs between training and serving environments.

import joblib
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_iris

X, y = load_iris(return_X_y=True)
pipe = Pipeline([('sc', StandardScaler()), ('lr', LogisticRegression())])
pipe.fit(X, y)

# Save and reload
joblib.dump(pipe, '/tmp/iris_pipeline.pkl')
loaded_pipe = joblib.load('/tmp/iris_pipeline.pkl')

print('Predictions match:', (pipe.predict(X) == loaded_pipe.predict(X)).all())

Pipeline Cross-Validation Best Practice

Always wrap your Pipeline in cross_val_score rather than manually looping over folds. This ensures that the scaler is re-fitted on each training fold and never touches the validation fold, giving you an honest estimate of generalisation performance.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC
from sklearn.datasets import load_digits
from sklearn.model_selection import cross_val_score
import numpy as np

X, y = load_digits(return_X_y=True)

pipe = Pipeline([
    ('sc', StandardScaler()),
    ('svm', SVC(kernel='rbf', C=5.0, gamma='scale'))
])

scores = cross_val_score(pipe, X, y, cv=5, n_jobs=-1)
print(f'CV accuracy: {np.mean(scores):.4f} +/- {np.std(scores):.4f}')

Quick Check

Test your understanding of scikit-learn Pipelines from this lesson.

Lesson Recap

In this lesson you learned: a Pipeline chains steps into one estimator, preventing leakage by fitting each step only on the training data it sees, make_pipeline provides auto-named shortcuts while explicit names improve grid-search readability, and a fitted Pipeline can be pickled and shared as a single artefact for consistent preprocessing at serving time. Next up we add ColumnTransformer inside a Pipeline to handle mixed numeric and categorical data.

無料で開始

AI チューターと学ぶ Python — 無料

ブラウザでリアルコードを書いて実行し、24/7 の AI チューターから瞬時にサポートを受け、ウェブまたはアプリで続きから学習できます。

コース
30
レッスン
120

よくある質問

「最初のPipelineを構築する:スケーラーと分類器」レッスンは無料ですか?

はい。「最初のPipelineを構築する:スケーラーと分類器」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Machine Learning Academyコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Machine Learning Academyコースには全4レッスンが含まれています。

「最初のPipelineを構築する:スケーラーと分類器」で何を学びますか?

StandardScalerとLogisticRegressionをPipelineに連結し、fitとpredictを呼び出して、スケーラーが学習データだけでfitされていることを確認します。 ブラウザで直接実行するハンズオンコードでMachine Learning Academyを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。

Machine Learning Academyを始めるのに経験は必要ですか?

事前経験は必要ありません。CoddyKitのMachine Learning Academyは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン1/4です。

「最初のPipelineを構築する:スケーラーと分類器」レッスンにはどのくらい時間がかかりますか?

ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。

このMachine Learning Academyレッスンでコードを書いて実行できますか?

はい。すべてのMachine Learning Academyレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。

このコースのすべてのレッスン

  1. 最初のPipelineを構築する:スケーラーと分類器
  2. Pipeline内のColumnTransformer
  3. 完全なPipelineの交差検証とグリッドサーチ
  4. joblibによるPipelineの保存と読み込み
← Machine Learning Academyに戻る