0Pricing
Machine Learning Academy · レッスン

SHAP値:グローバルおよびローカルな特徴量重要度

勾配ブースティングモデルのSHAP値を計算し、beeswarmプロットと棒グラフによる要約を作成して、技術者ではない関係者に単一の予測結果を説明します。

「SHAP値:グローバルおよびローカルな特徴量重要度」はCoddyKit上の無料Machine Learning Academyレッスンです。 これはレッスン1/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはMachine Learning Academy学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Machine Learning Academyコースには全4レッスンが含まれています。

このレッスンの一部はまだ翻訳されておらず、英語で表示されています。

Why Model Explainability Matters

Explainability is the ability to understand why a model made a specific prediction. In high-stakes domains like lending, healthcare, and hiring, regulators and users demand explanations — not just accurate predictions. SHAP (SHapley Additive exPlanations) provides a mathematically principled framework rooted in cooperative game theory to deliver these explanations for any model.

Shapley Values: The Game Theory Origin

SHAP values borrow from Shapley values in cooperative game theory, where players (features) collaborate to produce an outcome (prediction). Each feature receives a fair share of credit by averaging its marginal contribution across all possible feature orderings. This makes SHAP the only additive attribution method satisfying the axioms of efficiency, symmetry, dummy, and additivity.

Installing and Importing SHAP

The shap library supports tree models, neural networks, and any black-box model. Install it with pip install shap and import it alongside your trained model. SHAP's explainers are model-type-aware: TreeExplainer for gradient boosting and random forests gives exact values in O(T·D²) time, far faster than the naive exponential-time Shapley computation.

import shap
import xgboost as xgb
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split

data = load_breast_cancer()
X_train, X_test, y_train, y_test = train_test_split(
    data.data, data.target, test_size=0.2, random_state=42
)

model = xgb.XGBClassifier(n_estimators=100, use_label_encoder=False, eval_metric='logloss')
model.fit(X_train, y_train)

explainer = shap.TreeExplainer(model)

Computing SHAP Values for the Test Set

Call explainer.shap_values(X_test) to produce a matrix where each row is a sample and each column is a feature. The SHAP value for feature j in sample i represents the contribution of feature j to pushing the prediction away from the expected base value. Positive values push toward the positive class; negative values push toward the negative class.

shap_values = explainer.shap_values(X_test)

print('SHAP values shape:', shap_values.shape)  # (n_samples, n_features)
print('Base value (expected prediction):', explainer.expected_value)
print('First sample SHAP values:', shap_values[0])

Global Importance: Bar Plot

Global feature importance summarises which features matter most across all predictions. The SHAP summary_plot in bar mode shows the mean absolute SHAP value per feature, ranking them from most to least important. This replaces the naive built-in feature importance that only counts split counts, which is biased toward high-cardinality features.

import matplotlib.pyplot as plt

# Bar plot: mean |SHAP| per feature
shap.summary_plot(shap_values, X_test,
                  feature_names=data.feature_names,
                  plot_type='bar')
plt.tight_layout()
plt.savefig('shap_bar.png', dpi=150)

Global Importance: Beeswarm Plot

The beeswarm plot (default summary_plot) is richer than a bar chart: each dot represents one sample, coloured by feature value (red = high, blue = low). The x-axis shows the SHAP value, so you can see not only which features matter but also in which direction a high or low feature value pushes predictions. This reveals nonlinear and interaction effects at a glance.

shap.summary_plot(shap_values, X_test,
                  feature_names=data.feature_names)
# Dots to the right = positive contribution to predicted class
# Red dots far right = high feature value strongly increases prediction

Local Explanation: Force Plot

A force plot explains a single prediction. It shows the base value on the left and the final prediction on the right, with features as arrows that push the output higher (red) or lower (blue). The width of each arrow is proportional to the feature's SHAP value. This is the explanation you would show a loan officer asking 'why was this application denied?'

# Explain the first test sample
i = 0
shap.force_plot(
    explainer.expected_value,
    shap_values[i],
    X_test[i],
    feature_names=data.feature_names,
    matplotlib=True
)

Local Explanation: Waterfall Plot

The waterfall plot is a cleaner alternative to the force plot for a single sample. It stacks SHAP contributions vertically from the base value, showing each feature's contribution as a bar segment. Positive contributions are red and push toward the top; negative contributions are blue and pull down. The final stack total equals the model's raw output for that sample.

import shap

explanation = shap.Explanation(
    values=shap_values[0],
    base_values=explainer.expected_value,
    data=X_test[0],
    feature_names=list(data.feature_names)
)
shap.waterfall_plot(explanation)

Dependence Plot: Feature Interactions

A SHAP dependence plot shows how a single feature's SHAP value changes as its raw value changes, coloured by a second feature to reveal interactions. For example, plotting 'worst radius' coloured by 'mean texture' reveals whether the effect of radius depends on texture. This goes beyond ordinary partial-dependence plots by accounting for all feature interactions naturally.

shap.dependence_plot(
    'worst radius',          # feature to plot on x-axis
    shap_values,
    X_test,
    feature_names=list(data.feature_names),
    interaction_index='mean texture'  # colour by this feature
)

SHAP with Any Model: KernelExplainer

When the model is a black box (SVM, neural network, any sklearn estimator), use shap.KernelExplainer, which approximates Shapley values by sampling coalitions and fitting a weighted linear model locally. It is model-agnostic but slower than TreeExplainer. Provide a background dataset summary (e.g., K-Means centroids) to speed up computation on large datasets.

from sklearn.svm import SVC
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
import shap
import numpy as np

pipeline = Pipeline([('scaler', StandardScaler()), ('svm', SVC(probability=True))])
pipeline.fit(X_train, y_train)

# Use 50 background samples for speed
background = shap.kmeans(X_train, 50)
explainer_k = shap.KernelExplainer(pipeline.predict_proba, background)
shap_vals_k = explainer_k.shap_values(X_test[:10])  # Explain 10 samples

Validation: SHAP Values Sum to Prediction

A key property of SHAP is efficiency: the sum of all SHAP values for a sample plus the base value must equal the model's raw output. Verifying this sanity check confirms the explainer is working correctly. Any discrepancy indicates a mismatch between the explainer type and the model, or incorrect background data.

import numpy as np

# For tree models, verify SHAP values sum to log-odds output
base = explainer.expected_value
for i in range(5):
    shap_sum = shap_values[i].sum() + base
    raw_pred = model.predict(X_test[i:i+1], output_margin=True)[0]
    print(f'Sample {i}: SHAP sum={shap_sum:.4f}, model output={raw_pred:.4f}, match={abs(shap_sum-raw_pred)<1e-4}')

Quick Check

Test your understanding of Machine Learning with Python concepts from this lesson.

Lesson Recap

In this lesson you learned: SHAP values quantify each feature's contribution to a prediction using Shapley values from game theory, global summaries (bar and beeswarm plots) reveal overall feature importance and direction, and local explanations (force and waterfall plots) justify individual predictions. Next up we explore LIME as an alternative model-agnostic explanation approach.

よくある質問

「SHAP値:グローバルおよびローカルな特徴量重要度」レッスンは無料ですか?

はい。「SHAP値:グローバルおよびローカルな特徴量重要度」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Machine Learning Academyコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Machine Learning Academyコースには全4レッスンが含まれています。

「SHAP値:グローバルおよびローカルな特徴量重要度」で何を学びますか?

勾配ブースティングモデルのSHAP値を計算し、beeswarmプロットと棒グラフによる要約を作成して、技術者ではない関係者に単一の予測結果を説明します。 ブラウザで直接実行するハンズオンコードでMachine Learning Academyを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。

Machine Learning Academyを始めるのに経験は必要ですか?

事前経験は必要ありません。CoddyKitのMachine Learning Academyは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン1/4です。

「SHAP値:グローバルおよびローカルな特徴量重要度」レッスンにはどのくらい時間がかかりますか?

ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。

このMachine Learning Academyレッスンでコードを書いて実行できますか?

はい。すべてのMachine Learning Academyレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。

このコースのすべてのレッスン

  1. SHAP値:グローバルおよびローカルな特徴量重要度
  2. LIME:ローカルな解釈可能モデル非依存説明
  3. 公平性指標:人口統計的パリティと均等な機会
  4. バイアス軽減戦略:前処理・処理中・後処理
← Machine Learning Academyに戻る