0Pricing
Machine Learning Academy · 课时

模型注册表:暂存、生产与归档

您将在 MLflow Model Registry 中注册模型版本,使其从 Staging 过渡到 Production,并使用 Python API 编写晋级工作流。

模型注册表:暂存、生产与归档 是 CoddyKit 上的免费 Machine Learning Academy 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Machine Learning Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Machine Learning Academy 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

What Is a Model Registry?

A model registry is a centralised catalogue that stores versioned trained models with their metadata. Instead of managing model files scattered across file systems, a registry provides a single source of truth with named versions, lifecycle stages (Staging, Production, Archived), and searchable annotations. The MLflow Model Registry is the most widely used open-source solution and integrates directly with the MLflow tracking server.

Registering a Model from a Run

After training, register the model by linking it to an existing MLflow run artifact. You can register directly during logging using the registered_model_name argument, or after the fact using the MLflow client. The registry creates a named model entry (e.g., 'SentimentClassifier') and assigns it Version 1. Subsequent registrations of the same model name automatically increment the version number.

import mlflow
import mlflow.sklearn
from sklearn.ensemble import RandomForestClassifier

# Option 1: Register during logging
with mlflow.start_run():
    clf = RandomForestClassifier(n_estimators=100, random_state=42)
    # clf.fit(X_train, y_train)
    mlflow.sklearn.log_model(
        sk_model=clf,
        artifact_path='model',
        registered_model_name='SentimentClassifier'  # auto-registers
    )
    print('Model registered as SentimentClassifier v1')

Using the MLflow Client for Registry Operations

The MlflowClient Python API gives programmatic control over the registry. Use it to register models from existing run artifacts, transition stages, and add descriptions — all from scripts rather than the UI. This is essential for automated CI/CD pipelines where a new model should be promoted only after passing evaluation tests, without requiring manual UI interaction from a data scientist.

from mlflow.tracking import MlflowClient

client = MlflowClient(tracking_uri='http://localhost:5000')

# Option 2: Register from an existing run artifact
run_id = 'abc123def456'  # get this from mlflow.last_active_run().info.run_id
model_uri = f'runs:/{run_id}/model'

model_version = mlflow.register_model(
    model_uri=model_uri,
    name='SentimentClassifier'
)
print('Version:', model_version.version)
print('Status:', model_version.status)  # PENDING_REGISTRATION -> READY

Lifecycle Stages: None, Staging, Production, Archived

Every model version in the registry has a lifecycle stage. New versions start at None. After automated evaluation passes, promote to Staging for integration testing. After Staging passes, promote to Production — the version serving live traffic. When a newer version supersedes it, move it to Archived to preserve history without deleting it. Only one version should be in Production at a time per model name.

from mlflow.tracking import MlflowClient

client = MlflowClient()

# Transition version 1 to Staging
client.transition_model_version_stage(
    name='SentimentClassifier',
    version='1',
    stage='Staging',
    archive_existing_versions=False
)
print('Version 1 -> Staging')

# After testing, promote to Production (archives previous Production)
client.transition_model_version_stage(
    name='SentimentClassifier',
    version='1',
    stage='Production',
    archive_existing_versions=True  # auto-archives old Production
)
print('Version 1 -> Production')

Adding Descriptions and Tags to Versions

Model versions should carry human-readable metadata. Add a description explaining what changed in this version: training data, preprocessing, or algorithm. Add tags for quick filtering, such as the deployment environment or dataset version. Good metadata makes it possible to answer audit questions ('What model was serving in February?') months after deployment without digging through git history.

from mlflow.tracking import MlflowClient

client = MlflowClient()

# Add description to version
client.update_model_version(
    name='SentimentClassifier',
    version='1',
    description=('RandomForest trained on IMDB v2 (50k reviews). '
                 'Test accuracy 0.924, F1 0.921. '
                 'Replaces rule-based baseline.')
)

# Add tags for filtering and search
client.set_model_version_tag(
    name='SentimentClassifier',
    version='1',
    key='dataset',
    value='imdb_v2'
)
client.set_model_version_tag('SentimentClassifier', '1', 'algorithm', 'random_forest')
print('Description and tags added.')

Loading the Production Model for Inference

In your inference service, always load the model by stage alias ('Production') rather than a hardcoded version number. This way, when you promote a new version to Production, the inference service automatically uses the new model on the next load without code changes. The models:/ URI scheme is a powerful MLflow convention for stage-based loading.

import mlflow.sklearn

# Load the current Production model by stage
model_name = 'SentimentClassifier'
stage = 'Production'
model_uri = f'models:/{model_name}/{stage}'

production_model = mlflow.sklearn.load_model(model_uri)
print('Loaded model from:', model_uri)

# Or load a specific version
version_uri = f'models:/{model_name}/1'
v1_model = mlflow.sklearn.load_model(version_uri)
print('Loaded specific version 1')

# Make predictions
# predictions = production_model.predict(X_new)

Searching and Comparing Versions

As more versions accumulate, use the client's search methods to filter by stage, tags, or metrics. Compare version performance programmatically: fetch the run ID associated with each version, query the run's metrics, and find the best-performing version to promote. This automation prevents manual errors and ensures promotion decisions are based on objective metric comparisons rather than guesswork.

from mlflow.tracking import MlflowClient

client = MlflowClient()

# List all versions of a model
versions = client.search_model_versions("name='SentimentClassifier'")
for v in versions:
    print(f'Version {v.version}: stage={v.current_stage}, run_id={v.run_id[:8]}')

# Get the metric from the associated training run
for v in versions:
    run = client.get_run(v.run_id)
    acc = run.data.metrics.get('test_accuracy', 'N/A')
    print(f'  Version {v.version} accuracy: {acc}')

Automated Promotion Script

A retraining pipeline should automatically promote a new model to Staging only if it outperforms the current Production model on a held-out evaluation set. This champion/challenger pattern prevents regression: the Production model is the champion, and the new model is the challenger. The challenger is promoted only if it beats the champion on the agreed metric (e.g., F1 on the validation set).

from mlflow.tracking import MlflowClient
import mlflow.sklearn

client = MlflowClient()

def get_metric(run_id, metric_name):
    return client.get_run(run_id).data.metrics.get(metric_name, 0)

def promote_if_better(model_name, challenger_version, metric='test_f1'):
    # Get current production version
    prod_versions = client.get_latest_versions(model_name, stages=['Production'])
    if not prod_versions:
        print('No production model found -- promoting challenger directly.')
        client.transition_model_version_stage(model_name, challenger_version, 'Production')
        return

    prod_v = prod_versions[0]
    prod_score = get_metric(prod_v.run_id, metric)
    chall_run_id = client.get_model_version(model_name, challenger_version).run_id
    chall_score = get_metric(chall_run_id, metric)

    print(f'Champion {metric}: {prod_score:.4f}  Challenger: {chall_score:.4f}')
    if chall_score > prod_score:
        client.transition_model_version_stage(model_name, challenger_version,
                                              'Production', archive_existing_versions=True)
        print('Challenger promoted to Production!')
    else:
        print('Champion retained.')

Archiving Superseded Models

When a new version enters Production, old Production versions should move to Archived rather than being deleted. Archived models are excluded from get_latest_versions queries but remain downloadable for audit, rollback, or future comparison. Never delete model versions in a regulated industry: financial services and healthcare require full version history for compliance audits.

from mlflow.tracking import MlflowClient

client = MlflowClient()

# Manually archive a specific version
client.transition_model_version_stage(
    name='SentimentClassifier',
    version='1',
    stage='Archived'
)
print('Version 1 archived.')

# List only archived versions
archived = client.search_model_versions(
    "name='SentimentClassifier' and stage='Archived'"
)
for v in archived:
    print(f'Archived: v{v.version} created {v.creation_timestamp}')

Model Serving with mlflow models serve

MLflow can serve any registered model as a local REST API with a single command. The endpoint accepts JSON payloads and returns predictions. This is useful for rapid prototyping and integration testing before deploying to a cloud platform. For production, use container-based serving (Docker + FastAPI or MLflow's Docker export) for better scalability and monitoring.

# Serve the Production model as a local REST endpoint
# mlflow models serve -m 'models:/SentimentClassifier/Production' --port 8080

# Then call it with curl:
# curl -X POST http://localhost:8080/invocations \
#   -H 'Content-Type: application/json' \
#   -d '{"dataframe_records": [{"feature1": 0.5, "feature2": 1.2}]}'

# Or with Python requests:
import requests
data = {'dataframe_records': [{'feature1': 0.5, 'feature2': 1.2}]}
response = requests.post('http://localhost:8080/invocations', json=data)
print('Prediction:', response.json())

Registry Webhooks and Notifications

MLflow Registry (in Databricks and some enterprise setups) supports webhooks that fire HTTP callbacks when model transitions occur. For open-source MLflow, simulate webhooks by polling the registry in a cron job. Common automation patterns include: sending a Slack notification when a model enters Staging, triggering integration tests when a model reaches Staging, and alerting the team when Production is updated.

# Polling script (run on a schedule, e.g., cron every 5 minutes)
from mlflow.tracking import MlflowClient
import json
import os

client = MlflowClient()
state_file = '/tmp/model_registry_state.json'

def load_state():
    if os.path.exists(state_file):
        return json.load(open(state_file))
    return {}

def save_state(state):
    json.dump(state, open(state_file, 'w'))

state = load_state()
prod = client.get_latest_versions('SentimentClassifier', stages=['Production'])
if prod:
    current_prod = prod[0].version
    if state.get('production_version') != current_prod:
        print(f'ALERT: Production changed to version {current_prod}')
        # send_slack_notification(current_prod)
        state['production_version'] = current_prod
        save_state(state)

Quick Check

Test your understanding of Machine Learning with Python concepts from this lesson.

Lesson Recap

In this lesson you learned: the MLflow Model Registry provides versioned model storage with lifecycle stages: None, Staging, Production, and Archived, load models by stage alias ('Production') rather than version number to enable seamless updates without code changes, and automated promotion scripts implement the champion/challenger pattern to prevent production regressions. Next up we build a GitHub Actions workflow that automatically retrains and promotes a model when new data arrives.

常见问题解答

「模型注册表:暂存、生产与归档」课时是免费的吗?

是的 — 「模型注册表:暂存、生产与归档」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Machine Learning Academy 课程的其余内容,请升级到 CoddyKit PRO。 Machine Learning Academy 课程共包含 4 节课。

「模型注册表:暂存、生产与归档」这节课中我会学到什么?

您将在 MLflow Model Registry 中注册模型版本,使其从 Staging 过渡到 Production,并使用 Python API 编写晋级工作流。 你通过在浏览器中直接运行的动手代码来练习 Machine Learning Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Machine Learning Academy 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Machine Learning Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。

「模型注册表:暂存、生产与归档」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Machine Learning Academy 课中编写并运行代码吗?

能。每节 Machine Learning Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 使用 MLflow 跟踪实验:记录参数、指标与制品
  2. 使用 Docker 为机器学习构建可复现环境
  3. 模型注册表:暂存、生产与归档
  4. 使用 GitHub Actions 实现自动重新训练管道
← 返回 Machine Learning Academy