0Pricing
Machine Learning Academy · レッスン

joblibとpickleによるモデルの保存

学習済みパイプラインをjoblibとpickleの両方でシリアライズして読み戻し、予測が一致することを確認して、正常に往復保存できたことを検証します。

「joblibとpickleによるモデルの保存」はCoddyKit上の無料Machine Learning Academyレッスンです。 これはレッスン1/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはMachine Learning Academy学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Machine Learning Academyコースには全4レッスンが含まれています。

このレッスンの一部はまだ翻訳されておらず、英語で表示されています。

Why Model Persistence Matters

Training a machine learning model is expensive: it can take minutes to hours and consumes significant compute. Model persistence saves the fitted model to disk so you can reload it instantly for inference without retraining. This is the bridge between the data science notebook and a production system — the serialised model file is the deployable artefact that data engineers package and serve.

What Gets Saved in a Model File?

When you serialise a fitted sklearn model or pipeline, the file captures: all fitted parameters (e.g., scaler mean and variance, tree structure, logistic regression coefficients), hyperparameter settings, and the Python class definition reference. It does NOT include the training data. Loading the file reconstructs a Python object ready to call predict immediately.

Saving with joblib.dump

joblib is the recommended serialisation tool for sklearn objects. It handles large NumPy arrays efficiently using memory mapping and supports transparent compression. The standard workflow is: train the model, dump it to a .joblib file, then load it in a separate script or service for inference.

import joblib
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_breast_cancer

X, y = load_breast_cancer(return_X_y=True)

pipe = Pipeline([
    ('scaler', StandardScaler()),
    ('clf', LogisticRegression(C=1.0, max_iter=300))
])
pipe.fit(X, y)

# Save
joblib.dump(pipe, '/tmp/cancer_model.joblib')
print('Model saved to /tmp/cancer_model.joblib')

Loading with joblib.load

joblib.load deserialises the file and returns the exact fitted pipeline object. The loaded model has all of the same attributes — named_steps, fitted scaler parameters, classifier coefficients — as the original. You can immediately call predict, predict_proba, or score without any additional setup.

import joblib
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split

X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0)

# Load in a fresh context
model = joblib.load('/tmp/cancer_model.joblib')

predictions = model.predict(X_test[:5])
print('Predictions:', predictions)
print('Test accuracy:', model.score(X_test, y_test).round(4))

Using pickle for Serialisation

Python's standard library pickle module also serialises sklearn objects. Open files in binary mode ('wb' for write, 'rb' for read). The pickle.HIGHEST_PROTOCOL constant uses the most efficient available protocol. For small models or scripting contexts, pickle is perfectly adequate.

import pickle
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_iris

X, y = load_iris(return_X_y=True)
model = LogisticRegression().fit(X, y)

# Save
with open('/tmp/iris_model.pkl', 'wb') as f:
    pickle.dump(model, f, protocol=pickle.HIGHEST_PROTOCOL)

# Load
with open('/tmp/iris_model.pkl', 'rb') as f:
    loaded = pickle.load(f)

print('Score:', loaded.score(X, y).round(4))
print('Coefficients shape:', loaded.coef_.shape)

Comparing joblib vs pickle File Sizes

For a large model like a RandomForest with 1000 trees, joblib's memory-mapped NumPy array storage is more efficient. The difference becomes especially pronounced when the model contains large parameter matrices. For small models (LogReg, SVM), the size difference is negligible.

import joblib
import pickle
import os
from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import make_classification

X, y = make_classification(n_samples=1000, n_features=20, random_state=0)
rf = RandomForestClassifier(n_estimators=100, random_state=0).fit(X, y)

# joblib
joblib.dump(rf, '/tmp/rf_model.joblib')

# pickle
with open('/tmp/rf_model.pkl', 'wb') as f:
    pickle.dump(rf, f)

print(f'joblib size: {os.path.getsize("/tmp/rf_model.joblib"):,} bytes')
print(f'pickle size: {os.path.getsize("/tmp/rf_model.pkl"):,} bytes')

Compression with joblib

Use joblib.dump(model, path, compress=3) to compress the file using zlib. Compression levels range from 1 (fast, larger) to 9 (slow, smallest). Level 3 is a practical default. For LZ4 compression (faster than zlib): compress=('lz4', 1). Load time slightly increases for compressed files but the network transfer and storage savings are often worth it.

import joblib
import os
from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import make_classification

X, y = make_classification(n_samples=1000, random_state=0)
rf = RandomForestClassifier(n_estimators=100, random_state=0).fit(X, y)

for level in [0, 3, 6, 9]:
    path = f'/tmp/rf_compress_{level}.joblib'
    joblib.dump(rf, path, compress=level)
    size = os.path.getsize(path)
    print(f'compress={level}: {size:,} bytes')

Verifying Round-Trip Consistency

After loading, always verify that the loaded model produces identical predictions to the original. This guards against silent corruption, version mismatches, or incomplete file writes. Compare predictions with np.array_equal on the same input.

import joblib
import numpy as np
from sklearn.datasets import load_iris
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

X, y = load_iris(return_X_y=True)
pipe = Pipeline([('sc', StandardScaler()), ('lr', LogisticRegression())]).fit(X, y)
original_preds = pipe.predict(X)

joblib.dump(pipe, '/tmp/verify_pipe.joblib')
loaded = joblib.load('/tmp/verify_pipe.joblib')
loaded_preds = loaded.predict(X)

if np.array_equal(original_preds, loaded_preds):
    print('Round-trip PASSED: predictions are identical.')
else:
    diff = (original_preds != loaded_preds).sum()
    print(f'Round-trip FAILED: {diff} different predictions!')

Version Pinning and Metadata

A model pickled with scikit-learn 1.2 may not load cleanly in scikit-learn 1.5 due to internal changes. Always store a metadata file alongside the model that records: sklearn version, Python version, training date, dataset version, and key metrics. This is your model card — the governance document that explains what the model is and how it was produced.

import json
import sklearn
import sys
from datetime import datetime

metadata = {
    'model_file': 'cancer_model.joblib',
    'sklearn_version': sklearn.__version__,
    'python_version': sys.version.split()[0],
    'training_date': datetime.utcnow().isoformat(),
    'dataset': 'breast_cancer',
    'test_accuracy': 0.9789,
    'features': 30,
    'algorithm': 'LogisticRegression'
}

with open('/tmp/cancer_model_metadata.json', 'w') as f:
    json.dump(metadata, f, indent=2)

print(json.dumps(metadata, indent=2))

Security: Never Load Untrusted Pickle Files

Critical security warning: both pickle and joblib can execute arbitrary Python code when loading. Never load a model file from an untrusted source — it could be a malicious payload disguised as a model. For models shared across organisations, consider safer formats: ONNX (Open Neural Network Exchange) is a standardised, inspectable format supported by most frameworks.

File Naming Conventions

Good naming conventions embed key information into the filename: algorithm, dataset, date, and metric. This makes the model registry self-documenting and prevents accidentally loading the wrong model version in production.

from datetime import date
from sklearn.metrics import accuracy_score
import joblib

# Example naming convention
dataset = 'breast_cancer'
algorithm = 'logreg'
test_acc = 0.9789
today = date.today().strftime('%Y%m%d')

filename = f'{dataset}_{algorithm}_{today}_acc{int(test_acc*100)}.joblib'
print('Model filename:', filename)
# e.g.: breast_cancer_logreg_20260620_acc97.joblib

# Load the model we saved earlier (demo)
model = joblib.load('/tmp/cancer_model.joblib')
print('Loaded OK')

Quick Check

Test your understanding of model serialisation with joblib and pickle from this lesson.

Lesson Recap

In this lesson you learned: joblib.dump and joblib.load serialise and restore fitted sklearn models efficiently, pickle works too but joblib is preferred for models with large NumPy arrays due to memory mapping, and always store a metadata file alongside the model recording library versions, training date, and key metrics for governance. Next up we design a versioning naming convention and metadata sidecar to track multiple model versions in a model registry.

よくある質問

「joblibとpickleによるモデルの保存」レッスンは無料ですか?

はい。「joblibとpickleによるモデルの保存」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Machine Learning Academyコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Machine Learning Academyコースには全4レッスンが含まれています。

「joblibとpickleによるモデルの保存」で何を学びますか?

学習済みパイプラインをjoblibとpickleの両方でシリアライズして読み戻し、予測が一致することを確認して、正常に往復保存できたことを検証します。 ブラウザで直接実行するハンズオンコードでMachine Learning Academyを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。

Machine Learning Academyを始めるのに経験は必要ですか?

事前経験は必要ありません。CoddyKitのMachine Learning Academyは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン1/4です。

「joblibとpickleによるモデルの保存」レッスンにはどのくらい時間がかかりますか?

ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。

このMachine Learning Academyレッスンでコードを書いて実行できますか?

はい。すべてのMachine Learning Academyレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。

このコースのすべてのレッスン

  1. joblibとpickleによるモデルの保存
  2. モデルのバージョン管理:ファイル名とメタデータが重要な理由
  3. FastAPIエンドポイントによる予測の提供
  4. 予測の監視:入力と出力のロギング
← Machine Learning Academyに戻る