0Pricing
Machine Learning Academy · 课时

机器学习工作流:从数据到预测

您将完整了解从原始数据收集和清洗,到模型训练、评估和部署的端到端流程

机器学习工作流:从数据到预测 是 CoddyKit 上的免费 Machine Learning Academy 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Machine Learning Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Machine Learning Academy 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

The End-to-End ML Pipeline

Building ML is more than training a model — it's a full pipeline. Knowing all the stages keeps you from rushing straight to modelling too soon.

Stage 1: Define the Problem

Stage one is to define the problem: what are you predicting, why does it matter, and how will you measure success? Vague goals make vague models.

Stage 2: Collect and Load Data

Next, collect and load your data from files, databases, or APIs — Pandas is the go-to tool. Always inspect the raw data before doing anything else.

import pandas as pd

# Load data from a CSV file
df = pd.read_csv('housing.csv')

# First inspection
print(df.shape)       # (rows, columns)
print(df.dtypes)      # data types per column
print(df.head())      # first 5 rows
print(df.describe())  # summary statistics

Stage 3: Exploratory Data Analysis

Exploratory Data Analysis is the detective work: plot distributions, check correlations, and hunt for outliers before you build any model.

import matplotlib.pyplot as plt
import seaborn as sns
import pandas as pd

df = pd.read_csv('housing.csv')

# Distribution of target variable
df['price'].hist(bins=50)
plt.title('House Price Distribution')
plt.show()

# Correlation heat map
sns.heatmap(df.corr(), annot=True, cmap='coolwarm')
plt.show()

# Check for missing values
print(df.isnull().sum())

Stage 4: Preprocess the Data

Preprocessing turns messy data into a clean feature matrix — filling gaps, scaling, and encoding. Golden rule: fit only on training data to avoid leakage.

from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
import pandas as pd
import numpy as np

df = pd.read_csv('housing.csv')
X = df.drop('price', axis=1)
y = df['price']

# Split first, then fit scaler ONLY on train
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)  # fit + transform on train
X_test_scaled = scaler.transform(X_test)         # transform only on test

Stage 5: Train the Model

Now train the model: in scikit-learn it's one fit() call. Start simple with a baseline like a linear model before reaching for anything complex.

from sklearn.linear_model import LinearRegression
from sklearn.ensemble import RandomForestRegressor

# Simple baseline first
linear_model = LinearRegression()
linear_model.fit(X_train_scaled, y_train)

# More complex model
rf_model = RandomForestRegressor(n_estimators=100, random_state=42)
rf_model.fit(X_train_scaled, y_train)

print('Linear model trained.')
print('Random Forest trained.')

Stage 6: Evaluate the Model

Evaluate on held-out test data the model never saw. Pick the right metric — MAE and RMSE for numbers, accuracy and F1 for categories.

from sklearn.metrics import mean_absolute_error, r2_score
import numpy as np

y_pred = linear_model.predict(X_test_scaled)

mae = mean_absolute_error(y_test, y_pred)
r2 = r2_score(y_test, y_pred)
rmse = np.sqrt(((y_test - y_pred) ** 2).mean())

print(f'MAE:  {mae:.2f}')
print(f'RMSE: {rmse:.2f}')
print(f'R2:   {r2:.3f}')

Stage 7: Iterate and Improve

The first model is rarely the best. ML is iterative: based on your results, gather more data, engineer features, or tune settings — then try again.

Stage 8: Deploy the Model

Deployment makes your trained model available to real users — often as a web API, a batch job, or an on-device model in an app. A notebook alone adds no value.

import joblib

# Save the trained model
joblib.dump(linear_model, 'house_price_model.pkl')
print('Model saved.')

# Later, load and predict in production
loaded_model = joblib.load('house_price_model.pkl')
prediction = loaded_model.predict(X_test_scaled[:1])
print(f'Prediction: ${prediction[0]:,.0f}')

Stage 9: Monitor and Retrain

Deployment isn't the finish line. Data shifts over time, so you need monitoring to catch silent performance drops and trigger retraining when needed.

The Workflow as a Scikit-learn Pipeline

A scikit-learn Pipeline chains preprocessing and modelling into one object. It prevents leakage and makes deployment cleaner — save and load just one thing.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LinearRegression

# Chain preprocessing and model in one object
pipeline = Pipeline([
    ('scaler', StandardScaler()),
    ('model', LinearRegression())
])

# fit() applies scaler then trains the model
pipeline.fit(X_train, y_train)

# predict() applies scaler then predicts
y_pred = pipeline.predict(X_test)
print('Pipeline prediction done.')

Quick Check

Test your understanding of Machine Learning with Python concepts from this lesson.

Lesson Recap

Great progress! The ML workflow runs from problem definition to monitoring, preprocessing fits on training data only, and deployment is the start, not the end.

常见问题解答

「机器学习工作流:从数据到预测」课时是免费的吗?

是的 — 「机器学习工作流:从数据到预测」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Machine Learning Academy 课程的其余内容,请升级到 CoddyKit PRO。 Machine Learning Academy 课程共包含 4 节课。

「机器学习工作流:从数据到预测」这节课中我会学到什么?

您将完整了解从原始数据收集和清洗,到模型训练、评估和部署的端到端流程 你通过在浏览器中直接运行的动手代码来练习 Machine Learning Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Machine Learning Academy 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Machine Learning Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。

「机器学习工作流:从数据到预测」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Machine Learning Academy 课中编写并运行代码吗?

能。每节 Machine Learning Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 传统编程与机器学习
  2. 监督学习、无监督学习与强化学习
  3. 机器学习工作流:从数据到预测
  4. 现实世界中的机器学习:用例与局限
← 返回 Machine Learning Academy