Ihre erste Pipeline erstellen: Skalierer plus Klassifikator
Lernende verketten StandardScaler und LogisticRegression in einer Pipeline, rufen fit und predict auf und bestätigen, dass der Skalierer nur mit Trainingsdaten angepasst wurde.
Ihre erste Pipeline erstellen: Skalierer plus Klassifikator ist eine kostenlose Machine Learning Academy-Lektion auf CoddyKit. Dies ist Lektion 1 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des Machine Learning Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der Machine Learning Academy-Kurs umfasst insgesamt 4 Lektionen.
Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.
What Is a scikit-learn Pipeline?
A Pipeline chains multiple processing steps into a single estimator object. Each step except the last must implement fit and transform; the final step only needs fit and predict. Calling pipeline.fit(X, y) runs all steps in sequence, and pipeline.predict(X) passes data through all transforms before classifying. This eliminates bookkeeping errors and prevents data leakage.
Why Pipelines Prevent Leakage
If you fit a StandardScaler on the full dataset and then split into train/test, the scaler has seen test-set statistics — this is data leakage. A Pipeline solves this automatically: when you pass a Pipeline to cross_val_score or GridSearchCV, the entire pipeline (including the scaler) is re-fitted from scratch on each training fold, so the test fold never influences the scaler parameters.
Constructing Your First Pipeline
Create a Pipeline by passing a list of (name, estimator) tuples. The names are arbitrary strings you choose — they are used to reference steps later (e.g., for grid-search hyperparameters). The most common first pipeline is a scaler followed by a classifier.
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
pipe = Pipeline([
('scaler', StandardScaler()),
('clf', LogisticRegression(C=1.0, max_iter=200))
])
print('Steps:', [name for name, _ in pipe.steps])
print('Named steps:', list(pipe.named_steps.keys()))Fitting and Predicting
After construction, use the Pipeline exactly like any sklearn estimator: fit on training data, predict on test data, score for accuracy. Internally, fit calls fit_transform on all intermediate steps and fit on the final estimator.
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
pipe = Pipeline([
('scaler', StandardScaler()),
('clf', LogisticRegression(max_iter=300))
])
pipe.fit(X_train, y_train)
print('Test accuracy:', pipe.score(X_test, y_test).round(4))
y_pred = pipe.predict(X_test)
print('Predictions[:5]:', y_pred[:5])Confirming Scaler Fitted Only on Train Data
After fitting the Pipeline on training data, you can inspect each step's fitted parameters. The scaler inside the pipeline will have mean_ and scale_ attributes computed from the training set only — not the full dataset. This confirms the pipeline is doing the right thing.
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
import numpy as np
X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0)
pipe = Pipeline([('sc', StandardScaler()), ('lr', LogisticRegression(max_iter=300))])
pipe.fit(X_train, y_train)
# Scaler mean from pipeline vs computed from X_train
print('Pipeline scaler mean[0]:', pipe.named_steps['sc'].mean_[0].round(4))
print('Direct train mean[0]: ', X_train[:, 0].mean().round(4))Accessing Intermediate Outputs
To get the transformed output of a specific step, use pipeline[:-1].transform(X) or access individual steps via pipeline.named_steps['step_name']. You can also call pipeline[:'step_name'] with Python slice notation to get a sub-pipeline up to and including that step. This is useful for debugging or inspecting what the data looks like after scaling.
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_iris
X, y = load_iris(return_X_y=True)
pipe = Pipeline([('sc', StandardScaler()), ('lr', LogisticRegression())])
pipe.fit(X, y)
# Get scaled output (all steps except the last)
X_scaled = pipe[:-1].transform(X)
print('After scaling — mean per feature:', X_scaled.mean(axis=0).round(4))
print('After scaling — std per feature:', X_scaled.std(axis=0).round(4))make_pipeline: A Shortcut
make_pipeline creates a Pipeline without requiring you to specify step names manually — it generates names from the class names in lowercase. This is convenient for quick experiments but less readable in production code where explicit names help identify steps in grid-search parameter strings.
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import MinMaxScaler
from sklearn.neighbors import KNeighborsClassifier
from sklearn.datasets import load_iris
from sklearn.model_selection import cross_val_score
import numpy as np
X, y = load_iris(return_X_y=True)
# Auto-names: 'minmaxscaler' and 'kneighborsclassifier'
pipe = make_pipeline(MinMaxScaler(), KNeighborsClassifier(n_neighbors=5))
print('Step names:', list(pipe.named_steps.keys()))
print('CV accuracy:', cross_val_score(pipe, X, y, cv=5).mean().round(4))Pipeline with Probability Predictions
If the final estimator supports predict_proba, the Pipeline exposes it too. This allows you to use a Pipeline with any function that expects probability outputs — like ROC-AUC scoring, calibration curves, or threshold tuning — without needing to manually apply preprocessing before calling the model.
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0)
pipe = Pipeline([('sc', StandardScaler()), ('lr', LogisticRegression(max_iter=300))])
pipe.fit(X_train, y_train)
proba = pipe.predict_proba(X_test)[:, 1]
print('ROC-AUC:', roc_auc_score(y_test, proba).round(4))Setting Parameters After Construction
You can update any step parameter after building the Pipeline using set_params(step__param=value) with the double-underscore separator. This is the same syntax used in GridSearchCV. It is useful when you want to experiment with different settings without rebuilding the entire pipeline from scratch.
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
pipe = Pipeline([('sc', StandardScaler()), ('lr', LogisticRegression())])
# Change C and max_iter after construction
pipe.set_params(lr__C=10.0, lr__max_iter=500)
print('C after set_params:', pipe.named_steps['lr'].C)Pickling and Sharing Pipelines
A fitted Pipeline is a single Python object that can be serialised with pickle or joblib. Sharing the pipeline as one file ensures that the exact same preprocessing steps — with the exact same scaler parameters — are applied at prediction time, eliminating consistency bugs between training and serving environments.
import joblib
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_iris
X, y = load_iris(return_X_y=True)
pipe = Pipeline([('sc', StandardScaler()), ('lr', LogisticRegression())])
pipe.fit(X, y)
# Save and reload
joblib.dump(pipe, '/tmp/iris_pipeline.pkl')
loaded_pipe = joblib.load('/tmp/iris_pipeline.pkl')
print('Predictions match:', (pipe.predict(X) == loaded_pipe.predict(X)).all())Pipeline Cross-Validation Best Practice
Always wrap your Pipeline in cross_val_score rather than manually looping over folds. This ensures that the scaler is re-fitted on each training fold and never touches the validation fold, giving you an honest estimate of generalisation performance.
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC
from sklearn.datasets import load_digits
from sklearn.model_selection import cross_val_score
import numpy as np
X, y = load_digits(return_X_y=True)
pipe = Pipeline([
('sc', StandardScaler()),
('svm', SVC(kernel='rbf', C=5.0, gamma='scale'))
])
scores = cross_val_score(pipe, X, y, cv=5, n_jobs=-1)
print(f'CV accuracy: {np.mean(scores):.4f} +/- {np.std(scores):.4f}')Quick Check
Test your understanding of scikit-learn Pipelines from this lesson.
Lesson Recap
In this lesson you learned: a Pipeline chains steps into one estimator, preventing leakage by fitting each step only on the training data it sees, make_pipeline provides auto-named shortcuts while explicit names improve grid-search readability, and a fitted Pipeline can be pickled and shared as a single artefact for consistent preprocessing at serving time. Next up we add ColumnTransformer inside a Pipeline to handle mixed numeric and categorical data.
Häufig gestellte Fragen
Ist die Lektion „Ihre erste Pipeline erstellen: Skalierer plus Klassifikator“ kostenlos?
Ja — der vollständige Text von „Ihre erste Pipeline erstellen: Skalierer plus Klassifikator“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des Machine Learning Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der Machine Learning Academy-Kurs umfasst insgesamt 4 Lektionen.
Was lerne ich in „Ihre erste Pipeline erstellen: Skalierer plus Klassifikator“?
Lernende verketten StandardScaler und LogisticRegression in einer Pipeline, rufen fit und predict auf und bestätigen, dass der Skalierer nur mit Trainingsdaten angepasst wurde. Du übst Machine Learning Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.
Brauche ich Erfahrung, um Machine Learning Academy zu starten?
Keine Vorkenntnisse erforderlich. Machine Learning Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 1 von 4.
Wie lange dauert die Lektion „Ihre erste Pipeline erstellen: Skalierer plus Klassifikator“?
Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.
Kann ich in dieser Machine Learning Academy-Lektion Code schreiben und ausführen?
Ja. Jede Machine Learning Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.
Alle Lektionen in diesem Kurs
- Ihre erste Pipeline erstellen: Skalierer plus Klassifikator
- ColumnTransformer in einer Pipeline
- Eine vollständige Pipeline per Kreuzvalidierung und Grid Search optimieren
- Eine Pipeline mit joblib speichern und laden