Machine Learning Academy · Lección

Creación de su primer pipeline: escalador más clasificador

Encadenará StandardScaler y LogisticRegression en un Pipeline, ejecutará fit y predict y comprobará que el escalador solo se ajustó con los datos de entrenamiento.

Lección 1 de 413 pasos

Creación de su primer pipeline: escalador más clasificador es una lección gratuita de Machine Learning Academy en CoddyKit. Esta es la lección 1 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Machine Learning Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Machine Learning Academy incluye 4 lecciones en total.

Partes de esta lección aún no han sido traducidas y se muestran en inglés.

What Is a scikit-learn Pipeline?

A Pipeline chains multiple processing steps into a single estimator object. Each step except the last must implement fit and transform; the final step only needs fit and predict. Calling pipeline.fit(X, y) runs all steps in sequence, and pipeline.predict(X) passes data through all transforms before classifying. This eliminates bookkeeping errors and prevents data leakage.

Why Pipelines Prevent Leakage

If you fit a StandardScaler on the full dataset and then split into train/test, the scaler has seen test-set statistics — this is data leakage. A Pipeline solves this automatically: when you pass a Pipeline to cross_val_score or GridSearchCV, the entire pipeline (including the scaler) is re-fitted from scratch on each training fold, so the test fold never influences the scaler parameters.

Constructing Your First Pipeline

Create a Pipeline by passing a list of (name, estimator) tuples. The names are arbitrary strings you choose — they are used to reference steps later (e.g., for grid-search hyperparameters). The most common first pipeline is a scaler followed by a classifier.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

pipe = Pipeline([
    ('scaler', StandardScaler()),
    ('clf', LogisticRegression(C=1.0, max_iter=200))
])

print('Steps:', [name for name, _ in pipe.steps])
print('Named steps:', list(pipe.named_steps.keys()))

Fitting and Predicting

After construction, use the Pipeline exactly like any sklearn estimator: fit on training data, predict on test data, score for accuracy. Internally, fit calls fit_transform on all intermediate steps and fit on the final estimator.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split

X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

pipe = Pipeline([
    ('scaler', StandardScaler()),
    ('clf', LogisticRegression(max_iter=300))
])

pipe.fit(X_train, y_train)
print('Test accuracy:', pipe.score(X_test, y_test).round(4))

y_pred = pipe.predict(X_test)
print('Predictions[:5]:', y_pred[:5])

Confirming Scaler Fitted Only on Train Data

After fitting the Pipeline on training data, you can inspect each step's fitted parameters. The scaler inside the pipeline will have mean_ and scale_ attributes computed from the training set only — not the full dataset. This confirms the pipeline is doing the right thing.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
import numpy as np

X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0)

pipe = Pipeline([('sc', StandardScaler()), ('lr', LogisticRegression(max_iter=300))])
pipe.fit(X_train, y_train)

# Scaler mean from pipeline vs computed from X_train
print('Pipeline scaler mean[0]:', pipe.named_steps['sc'].mean_[0].round(4))
print('Direct train mean[0]:   ', X_train[:, 0].mean().round(4))

Accessing Intermediate Outputs

To get the transformed output of a specific step, use pipeline[:-1].transform(X) or access individual steps via pipeline.named_steps['step_name']. You can also call pipeline[:'step_name'] with Python slice notation to get a sub-pipeline up to and including that step. This is useful for debugging or inspecting what the data looks like after scaling.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_iris

X, y = load_iris(return_X_y=True)

pipe = Pipeline([('sc', StandardScaler()), ('lr', LogisticRegression())])
pipe.fit(X, y)

# Get scaled output (all steps except the last)
X_scaled = pipe[:-1].transform(X)
print('After scaling — mean per feature:', X_scaled.mean(axis=0).round(4))
print('After scaling — std per feature:', X_scaled.std(axis=0).round(4))

make_pipeline: A Shortcut

make_pipeline creates a Pipeline without requiring you to specify step names manually — it generates names from the class names in lowercase. This is convenient for quick experiments but less readable in production code where explicit names help identify steps in grid-search parameter strings.

from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import MinMaxScaler
from sklearn.neighbors import KNeighborsClassifier
from sklearn.datasets import load_iris
from sklearn.model_selection import cross_val_score
import numpy as np

X, y = load_iris(return_X_y=True)

# Auto-names: 'minmaxscaler' and 'kneighborsclassifier'
pipe = make_pipeline(MinMaxScaler(), KNeighborsClassifier(n_neighbors=5))
print('Step names:', list(pipe.named_steps.keys()))
print('CV accuracy:', cross_val_score(pipe, X, y, cv=5).mean().round(4))

Pipeline with Probability Predictions

If the final estimator supports predict_proba, the Pipeline exposes it too. This allows you to use a Pipeline with any function that expects probability outputs — like ROC-AUC scoring, calibration curves, or threshold tuning — without needing to manually apply preprocessing before calling the model.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split

X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0)

pipe = Pipeline([('sc', StandardScaler()), ('lr', LogisticRegression(max_iter=300))])
pipe.fit(X_train, y_train)

proba = pipe.predict_proba(X_test)[:, 1]
print('ROC-AUC:', roc_auc_score(y_test, proba).round(4))

Setting Parameters After Construction

You can update any step parameter after building the Pipeline using set_params(step__param=value) with the double-underscore separator. This is the same syntax used in GridSearchCV. It is useful when you want to experiment with different settings without rebuilding the entire pipeline from scratch.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

pipe = Pipeline([('sc', StandardScaler()), ('lr', LogisticRegression())])

# Change C and max_iter after construction
pipe.set_params(lr__C=10.0, lr__max_iter=500)
print('C after set_params:', pipe.named_steps['lr'].C)

Pickling and Sharing Pipelines

A fitted Pipeline is a single Python object that can be serialised with pickle or joblib. Sharing the pipeline as one file ensures that the exact same preprocessing steps — with the exact same scaler parameters — are applied at prediction time, eliminating consistency bugs between training and serving environments.

import joblib
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_iris

X, y = load_iris(return_X_y=True)
pipe = Pipeline([('sc', StandardScaler()), ('lr', LogisticRegression())])
pipe.fit(X, y)

# Save and reload
joblib.dump(pipe, '/tmp/iris_pipeline.pkl')
loaded_pipe = joblib.load('/tmp/iris_pipeline.pkl')

print('Predictions match:', (pipe.predict(X) == loaded_pipe.predict(X)).all())

Pipeline Cross-Validation Best Practice

Always wrap your Pipeline in cross_val_score rather than manually looping over folds. This ensures that the scaler is re-fitted on each training fold and never touches the validation fold, giving you an honest estimate of generalisation performance.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC
from sklearn.datasets import load_digits
from sklearn.model_selection import cross_val_score
import numpy as np

X, y = load_digits(return_X_y=True)

pipe = Pipeline([
    ('sc', StandardScaler()),
    ('svm', SVC(kernel='rbf', C=5.0, gamma='scale'))
])

scores = cross_val_score(pipe, X, y, cv=5, n_jobs=-1)
print(f'CV accuracy: {np.mean(scores):.4f} +/- {np.std(scores):.4f}')

Quick Check

Test your understanding of scikit-learn Pipelines from this lesson.

Lesson Recap

In this lesson you learned: a Pipeline chains steps into one estimator, preventing leakage by fitting each step only on the training data it sees, make_pipeline provides auto-named shortcuts while explicit names improve grid-search readability, and a fitted Pipeline can be pickled and shared as a single artefact for consistent preprocessing at serving time. Next up we add ColumnTransformer inside a Pipeline to handle mixed numeric and categorical data.

Gratis para empezar

Aprende Python con un tutor de IA — gratis

Escribe y ejecuta código real en tu navegador, obtén ayuda instantánea de un tutor de IA disponible 24/7 y continúa donde lo dejaste en la web o en la aplicación.

Cursos
30
Lecciones
120

Preguntas frecuentes

¿La lección «Creación de su primer pipeline: escalador más clasificador» es gratis?

Sí — el texto completo de «Creación de su primer pipeline: escalador más clasificador» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Machine Learning Academy, actualiza a CoddyKit PRO. El curso de Machine Learning Academy incluye 4 lecciones en total.

¿Qué aprenderé en «Creación de su primer pipeline: escalador más clasificador»?

Encadenará StandardScaler y LogisticRegression en un Pipeline, ejecutará fit y predict y comprobará que el escalador solo se ajustó con los datos de entrenamiento. Practicas Machine Learning Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.

¿Necesito experiencia previa para empezar Machine Learning Academy?

No se requiere experiencia previa. Machine Learning Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 1 de 4.

¿Cuánto tiempo toma la lección «Creación de su primer pipeline: escalador más clasificador»?

La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.

¿Puedo escribir y ejecutar código en esta lección de Machine Learning Academy?

Sí. Cada lección de Machine Learning Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.

Todas las lecciones de este curso

  1. Creación de su primer pipeline: escalador más clasificador
  2. ColumnTransformer dentro de un pipeline
  3. Validación cruzada y búsqueda en cuadrícula de un pipeline completo
  4. Guardado y carga de un pipeline con joblib
← Volver a Machine Learning Academy