تقديم التنبؤات عبر نقطة نهاية FastAPI
سيغلّف المتعلمون نموذجًا محفوظًا باستخدام joblib في مسار POST عبر FastAPI يقبل حمولة JSON ويعيد تنبؤًا، ثم يختبرونه بطلب curl.
تقديم التنبؤات عبر نقطة نهاية FastAPI درس مجاني في Machine Learning Academy على CoddyKit. هذا هو الدرس 3 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Machine Learning Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Machine Learning Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
From Notebook to Production API
A Jupyter notebook is a great development environment but a terrible production serving system. The standard path from notebook to production is: train and save a model with joblib, wrap it in a REST API, and deploy that API as a containerised service. The API accepts raw feature values as JSON, preprocesses them through the fitted pipeline, and returns predictions in milliseconds.
Why FastAPI for ML Serving?
FastAPI is a modern Python web framework built on Pydantic and Starlette. It generates automatic interactive documentation (Swagger UI), validates request bodies with type hints, and handles async I/O efficiently. For ML serving, FastAPI is popular because it requires very little boilerplate, supports concurrent requests via async workers, and integrates naturally with Python data types used in sklearn and pandas.
Installing FastAPI and Uvicorn
FastAPI requires uvicorn as the ASGI server to run it. Install both with a single command. uvicorn is a high-performance async server that handles HTTP connections and passes requests to the FastAPI application. In production, you would typically run uvicorn behind an nginx reverse proxy with multiple worker processes.
# pip install fastapi uvicorn[standard]
# Verify installation
from fastapi import FastAPI
from pydantic import BaseModel
import uvicorn
print('FastAPI ready')Defining the Request Schema with Pydantic
Pydantic BaseModel classes define the structure of incoming requests. FastAPI uses these models to automatically validate JSON bodies — if a required field is missing or has the wrong type, FastAPI returns a clear 422 error before your code even runs. Each field in the Pydantic model corresponds to one input feature for the model.
from pydantic import BaseModel
from typing import Optional
class IrisFeatures(BaseModel):
sepal_length: float
sepal_width: float
petal_length: float
petal_width: float
class PredictionResponse(BaseModel):
predicted_class: int
class_name: str
confidence: float
# Example input (FastAPI will validate this automatically)
input_data = IrisFeatures(sepal_length=5.1, sepal_width=3.5,
petal_length=1.4, petal_width=0.2)
print('Input:', input_data)Loading the Model at Startup
Load the model once at startup, not on every request. Loading a joblib file on every prediction would add hundreds of milliseconds of latency per request. Use a module-level variable or a FastAPI lifespan event handler to load the model when the server starts and keep it in memory for all subsequent requests.
import joblib
from contextlib import asynccontextmanager
from fastapi import FastAPI
ml_models = {}
@asynccontextmanager
async def lifespan(app: FastAPI):
# Startup: load model once
ml_models['iris'] = joblib.load('/tmp/iris_pipeline.joblib')
print('Model loaded at startup')
yield
# Shutdown: cleanup if needed
ml_models.clear()
app = FastAPI(title='Iris Predictor API', lifespan=lifespan)Creating the Prediction Endpoint
Define a POST route that accepts the Pydantic input model, converts it to a NumPy array, calls pipeline.predict and predict_proba, and returns the prediction as a structured JSON response. FastAPI serialises Pydantic response models automatically.
import numpy as np
from fastapi import FastAPI
from pydantic import BaseModel
import joblib
app = FastAPI()
model = None
@app.on_event('startup')
def load_model():
global model
model = joblib.load('/tmp/iris_pipeline.joblib')
CLASS_NAMES = ['setosa', 'versicolor', 'virginica']
@app.post('/predict')
def predict(features: IrisFeatures):
X = np.array([[features.sepal_length, features.sepal_width,
features.petal_length, features.petal_width]])
pred = int(model.predict(X)[0])
proba = float(model.predict_proba(X)[0].max())
return {
'predicted_class': pred,
'class_name': CLASS_NAMES[pred],
'confidence': round(proba, 4)
}Adding a Health Check Endpoint
A /health or /ping endpoint is essential for production services. Load balancers and orchestration systems (Kubernetes, ECS) call this endpoint periodically to confirm the service is alive. A healthy response means the server is running AND the model is loaded. Return 503 if the model failed to load.
from fastapi import FastAPI
from fastapi.responses import JSONResponse
app = FastAPI()
@app.get('/health')
def health():
if model is None:
return JSONResponse(status_code=503,
content={'status': 'unhealthy', 'reason': 'model not loaded'})
return {'status': 'ok', 'model': 'iris_pipeline', 'version': '1.0.0'}
@app.get('/')
def root():
return {'message': 'Iris Predictor API — POST /predict to get a classification'}Running the Server Locally
Save the FastAPI app to main.py and start it with uvicorn main:app --reload. The --reload flag auto-restarts on file changes (development only). Navigate to http://localhost:8000/docs to see the auto-generated Swagger UI where you can test predictions interactively.
# Save to main.py then run:
# uvicorn main:app --host 0.0.0.0 --port 8000 --reload
# Test with curl:
# curl -X POST http://localhost:8000/predict \
# -H 'Content-Type: application/json' \
# -d '{"sepal_length": 5.1, "sepal_width": 3.5, "petal_length": 1.4, "petal_width": 0.2}'
#
# Expected response:
# {"predicted_class": 0, "class_name": "setosa", "confidence": 0.9981}
print('Command to start: uvicorn main:app --reload --port 8000')Testing the Endpoint with the requests Library
In a test script or notebook, use requests.post to call your running API. This is also how client applications (mobile apps, dashboards, other microservices) consume the prediction API. The same request format works from any language — curl, JavaScript fetch, Go's http.Client.
import requests
url = 'http://localhost:8000/predict'
payload = {
'sepal_length': 6.3,
'sepal_width': 3.3,
'petal_length': 6.0,
'petal_width': 2.5
}
response = requests.post(url, json=payload)
if response.status_code == 200:
result = response.json()
print('Predicted class:', result['class_name'])
print('Confidence:', result['confidence'])
else:
print('Error:', response.status_code, response.text)Input Validation and Error Handling
FastAPI's Pydantic validation catches type errors automatically, but you should also handle model-level errors (e.g., unexpected NaN values, out-of-range inputs). Use try/except inside the route function and return a 400 or 500 with a meaningful error message. Avoid leaking internal error details (stack traces) to API callers in production.
from fastapi import FastAPI, HTTPException
import numpy as np
app = FastAPI()
@app.post('/predict')
def predict(features: IrisFeatures):
try:
X = np.array([[features.sepal_length, features.sepal_width,
features.petal_length, features.petal_width]])
if np.any(np.isnan(X)) or np.any(X < 0):
raise HTTPException(status_code=400,
detail='Input contains invalid values (NaN or negative)')
pred = int(model.predict(X)[0])
proba = float(model.predict_proba(X)[0].max())
return {'predicted_class': pred, 'confidence': round(proba, 4)}
except HTTPException:
raise
except Exception as e:
raise HTTPException(status_code=500, detail='Internal prediction error')Batch Prediction Endpoint
For high-throughput use cases, add a batch endpoint that accepts a list of feature sets and returns a list of predictions in one API call. Batching reduces network overhead and allows the model to vectorise predictions efficiently (sklearn predict handles matrices).
from typing import List
from pydantic import BaseModel
import numpy as np
class BatchRequest(BaseModel):
instances: List[IrisFeatures]
@app.post('/predict/batch')
def predict_batch(batch: BatchRequest):
X = np.array([[f.sepal_length, f.sepal_width, f.petal_length, f.petal_width]
for f in batch.instances])
preds = model.predict(X).tolist()
probas = model.predict_proba(X).max(axis=1).tolist()
return {'predictions': [{'class': p, 'confidence': round(c, 4)}
for p, c in zip(preds, probas)]}Quick Check
Test your understanding of serving ML predictions with FastAPI from this lesson.
Lesson Recap
In this lesson you learned: FastAPI wraps a joblib-loaded pipeline into a typed REST endpoint with automatic JSON validation and Swagger documentation, load the model once at startup to avoid per-request disk I/O latency, and always include a /health endpoint so load balancers and orchestrators can verify the service is alive. Next up we add prediction logging to the API and discuss data drift and model retraining triggers.
تعلم Python مع معلم ذكاء اصطناعي — مجانًا
اكتب وقم بتشغيل أكوادك الفعلية في المتصفح، واحصل على مساعدة فورية من معلم ذكاء اصطناعي متاح 24/7، واستمر من حيث توقفت على الويب أو في التطبيق.
- الدورات
- 30
- الدروس
- 120
الأسئلة الشائعة
هل درس «تقديم التنبؤات عبر نقطة نهاية FastAPI» مجاني؟
نعم — نص درس «تقديم التنبؤات عبر نقطة نهاية FastAPI» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Machine Learning Academy، انتقل إلى CoddyKit PRO. تتضمن دورة Machine Learning Academy 4 دروس في المجموع.
ماذا ستتعلم في «تقديم التنبؤات عبر نقطة نهاية FastAPI»؟
سيغلّف المتعلمون نموذجًا محفوظًا باستخدام joblib في مسار POST عبر FastAPI يقبل حمولة JSON ويعيد تنبؤًا، ثم يختبرونه بطلب curl. تتمرن على Machine Learning Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ Machine Learning Academy؟
لا تُشترط خبرة سابقة. Machine Learning Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 3 من أصل 4.
كم من الوقت يستغرق درس «تقديم التنبؤات عبر نقطة نهاية FastAPI»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس Machine Learning Academy هذا؟
نعم. كل درس في Machine Learning Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- حفظ النماذج باستخدام joblib وpickle
- إدارة إصدارات النماذج: أهمية أسماء الملفات والبيانات الوصفية
- تقديم التنبؤات عبر نقطة نهاية FastAPI
- مراقبة التنبؤات: تسجيل المدخلات والمخرجات