Pipeline内のColumnTransformer
数値データとカテゴリデータが混在する前処理用のColumnTransformerをPipeline内にネストし、異種の生データを手動で分割せずにそのまま入力できるようにします。
「Pipeline内のColumnTransformer」はCoddyKit上の無料Machine Learning Academyレッスンです。 これはレッスン2/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはMachine Learning Academy学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Machine Learning Academyコースには全4レッスンが含まれています。
このレッスンの一部はまだ翻訳されておらず、英語で表示されています。
The Mixed-Data Problem
Real-world tabular datasets almost always contain a mix of numeric columns (age, salary, temperature) and categorical columns (city, product category, gender). Each type needs different preprocessing: numeric columns need scaling or imputation, while categorical columns need encoding. ColumnTransformer lets you apply different transformers to different subsets of columns in parallel, producing a single clean feature matrix.
ColumnTransformer: The Basic Structure
ColumnTransformer takes a list of (name, transformer, columns) triples. The columns can be a list of column names, a list of integer indices, a boolean mask, or a sklearn selector like make_column_selector. After transformation, the results from all transformers are horizontally concatenated.
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
ct = ColumnTransformer([
('num', StandardScaler(), ['age', 'salary']),
('cat', OneHotEncoder(handle_unknown='ignore'), ['city', 'education'])
])
print('Transformers:', [name for name, _, _ in ct.transformers])Working with a Real Mixed Dataset
Let us create a small mixed DataFrame and apply a ColumnTransformer to see the output shape and values. This illustrates how numeric and categorical outputs are concatenated into a single NumPy array that any sklearn estimator can consume.
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
df = pd.DataFrame({
'age': [25, 40, 35, 55],
'salary': [50000, 80000, 65000, 90000],
'city': ['NYC', 'LA', 'NYC', 'SF'],
'education': ['BSc', 'MSc', 'BSc', 'PhD']
})
ct = ColumnTransformer([
('num', StandardScaler(), ['age', 'salary']),
('cat', OneHotEncoder(handle_unknown='ignore', sparse_output=False), ['city', 'education'])
])
X_transformed = ct.fit_transform(df)
print('Output shape:', X_transformed.shape)
print('Output:\n', X_transformed.round(2))Automatic Column Selection
Instead of listing columns manually, use make_column_selector to automatically select columns by dtype. dtype_include=np.number selects all numeric columns; dtype_exclude=np.number selects non-numeric (categorical/string) columns. This is especially helpful for wide datasets where listing every column name would be impractical.
import numpy as np
import pandas as pd
from sklearn.compose import ColumnTransformer, make_column_selector
from sklearn.preprocessing import StandardScaler, OneHotEncoder
df = pd.DataFrame({
'age': [25, 40, 35], 'salary': [50000, 80000, 65000],
'city': ['NYC', 'LA', 'NYC']
})
ct = ColumnTransformer([
('num', StandardScaler(), make_column_selector(dtype_include=np.number)),
('cat', OneHotEncoder(), make_column_selector(dtype_exclude=np.number))
])
print(ct.fit_transform(df).shape)Nesting ColumnTransformer Inside a Pipeline
The real power comes from nesting ColumnTransformer as a preprocessing step inside a Pipeline. The full pipeline — preprocessing and classifier — becomes a single estimator. All sklearn tooling (cross_val_score, GridSearchCV, joblib serialisation) works on this combined object.
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer, make_column_selector
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.linear_model import LogisticRegression
import numpy as np
preprocessor = ColumnTransformer([
('num', StandardScaler(), make_column_selector(dtype_include=np.number)),
('cat', OneHotEncoder(handle_unknown='ignore'),
make_column_selector(dtype_exclude=np.number))
])
pipe = Pipeline([
('preprocessor', preprocessor),
('clf', LogisticRegression(max_iter=300))
])
print('Pipeline steps:', [name for name, _ in pipe.steps])End-to-End Example with Titanic Data
Let us apply a ColumnTransformer pipeline to a Titanic-style dataset. Numeric columns (Age, Fare) get median imputation then standard scaling; categorical columns (Sex, Embarked) get most-frequent imputation then one-hot encoding. The pipeline is then fitted and scored.
import pandas as pd
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
import numpy as np
# Simulated Titanic subset
df = pd.DataFrame({
'Age': [22, 38, None, 35, 28],
'Fare': [7.25, 71.83, 7.92, 53.1, 8.05],
'Sex': ['male', 'female', 'female', 'male', 'male'],
'Embarked': ['S', 'C', 'S', None, 'S'],
'Survived': [0, 1, 1, 1, 0]
})
X = df.drop('Survived', axis=1)
y = df['Survived']
num_pipe = Pipeline([('impute', SimpleImputer(strategy='median')),
('scale', StandardScaler())])
cat_pipe = Pipeline([('impute', SimpleImputer(strategy='most_frequent')),
('encode', OneHotEncoder(handle_unknown='ignore'))])
preprocessor = ColumnTransformer([
('num', num_pipe, ['Age', 'Fare']),
('cat', cat_pipe, ['Sex', 'Embarked'])
])
pipe = Pipeline([('prep', preprocessor), ('clf', LogisticRegression())])
pipe.fit(X, y)
print('Fitted successfully. Output features:', pipe['prep'].transform(X).shape[1])The remainder Parameter
By default, ColumnTransformer drops columns not explicitly listed. Set remainder='passthrough' to pass through any unlisted columns unchanged, or remainder=StandardScaler() to apply a transformer to them. This is useful when you have many numeric columns and only want to specially handle a few categorical ones.
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder
df = pd.DataFrame({
'age': [25, 40, 35],
'salary': [50000, 80000, 65000],
'city': ['NYC', 'LA', 'NYC']
})
# Only encode city; pass through numeric columns
ct = ColumnTransformer([
('cat', OneHotEncoder(), ['city'])
], remainder='passthrough')
print(ct.fit_transform(df))Getting Feature Names After Transformation
After fitting, call columntransformer.get_feature_names_out() to retrieve names for all output columns. OneHotEncoder contributes names like cat__city_NYC; StandardScaler produces names like num__age. These names are essential for interpreting feature importances in tree models trained on the transformed data.
import pandas as pd
import numpy as np
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
df = pd.DataFrame({
'age': [25, 40, 35],
'salary': [50000, 80000, 65000],
'city': ['NYC', 'LA', 'NYC']
})
ct = ColumnTransformer([
('num', StandardScaler(), ['age', 'salary']),
('cat', OneHotEncoder(sparse_output=False), ['city'])
])
ct.fit(df)
print('Output feature names:')
print(ct.get_feature_names_out())Grid-Searching Pipeline with ColumnTransformer
To tune parameters of steps inside a ColumnTransformer nested in a Pipeline, use the double-underscore chain: preprocessor__num__scale__with_std or preprocessor__cat__encode__drop. The naming convention is pipeline_step__ct_step__substep__param. It looks verbose but is completely consistent.
from sklearn.model_selection import GridSearchCV
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.linear_model import LogisticRegression
# (assuming preprocessor and pipe are defined as above)
param_grid = {
'clf__C': [0.1, 1.0, 10.0],
# 'prep__num__scale__with_std': [True, False]
}
grid = GridSearchCV(pipe, param_grid, cv=3)
# grid.fit(X, y) # would run on real data
print('Grid ready with params:', list(param_grid.keys()))ColumnTransformer with Imputation Substeps
For each column type, you can nest its own sub-Pipeline inside the ColumnTransformer to handle multiple sequential operations. The numeric sub-pipe imputes then scales; the categorical sub-pipe imputes (to handle missing strings) then encodes. This structure keeps all preprocessing logic in one auditable object.
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
num_subpipe = Pipeline([
('imputer', SimpleImputer(strategy='median')),
('scaler', StandardScaler())
])
cat_subpipe = Pipeline([
('imputer', SimpleImputer(strategy='most_frequent')),
('encoder', OneHotEncoder(handle_unknown='ignore', sparse_output=False))
])
print('Numeric sub-pipeline steps:', [s[0] for s in num_subpipe.steps])
print('Categorical sub-pipeline steps:', [s[0] for s in cat_subpipe.steps])Validating the Full Pipeline
After building a complex pipeline, run a quick sanity check: fit on a small synthetic dataset, confirm the output shape matches expectations, and verify that predictions are sensible. Catching shape errors early saves debugging time later.
import pandas as pd
import numpy as np
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
X = pd.DataFrame({
'num1': np.random.randn(100),
'num2': np.random.randn(100),
'cat1': np.random.choice(['A', 'B', 'C'], 100)
})
y = np.random.randint(0, 2, 100)
ct = ColumnTransformer([
('num', Pipeline([('imp', SimpleImputer()), ('sc', StandardScaler())]), ['num1', 'num2']),
('cat', Pipeline([('imp', SimpleImputer(strategy='most_frequent')),
('enc', OneHotEncoder(sparse_output=False))]), ['cat1'])
])
pipe = Pipeline([('prep', ct), ('clf', LogisticRegression())])
pipe.fit(X, y)
print('Accuracy:', pipe.score(X, y).round(4))
print('Transformed shape:', ct.transform(X).shape)Quick Check
Test your understanding of ColumnTransformer from this lesson.
Lesson Recap
In this lesson you learned: ColumnTransformer applies different transformers to different column subsets simultaneously, nesting ColumnTransformer inside a Pipeline creates one auditable, leak-proof object for mixed-type preprocessing, and double-underscore notation lets you tune any nested parameter via GridSearchCV. Next up we cross-validate and grid-search a full pipeline to find optimal hyperparameters without leakage.
よくある質問
「Pipeline内のColumnTransformer」レッスンは無料ですか?
はい。「Pipeline内のColumnTransformer」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Machine Learning Academyコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Machine Learning Academyコースには全4レッスンが含まれています。
「Pipeline内のColumnTransformer」で何を学びますか?
数値データとカテゴリデータが混在する前処理用のColumnTransformerをPipeline内にネストし、異種の生データを手動で分割せずにそのまま入力できるようにします。 ブラウザで直接実行するハンズオンコードでMachine Learning Academyを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。
Machine Learning Academyを始めるのに経験は必要ですか?
事前経験は必要ありません。CoddyKitのMachine Learning Academyは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン2/4です。
「Pipeline内のColumnTransformer」レッスンにはどのくらい時間がかかりますか?
ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。
このMachine Learning Academyレッスンでコードを書いて実行できますか?
はい。すべてのMachine Learning Academyレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。
このコースのすべてのレッスン
- 最初のPipelineを構築する:スケーラーと分類器
- Pipeline内のColumnTransformer
- 完全なPipelineの交差検証とグリッドサーチ
- joblibによるPipelineの保存と読み込み