使用 Prometheus 暴露指标
跟踪请求速率、错误和持续时间
使用 Prometheus 暴露指标 是 CoddyKit 上的免费 MLOps Academy 课时。 这是第 2 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 MLOps Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 MLOps Academy 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Logs vs Metrics
Logs tell the story of one request. Metrics are numbers aggregated over many requests, so you can see trends like rate and error count at a glance. 📈
Meet Prometheus
Prometheus is the standard tool for collecting these numbers. Your service exposes metrics, and Prometheus regularly reads them for you.
Pull, Not Push
Prometheus uses a pull model: it scrapes a metrics URL on your service every few seconds. Your job is just to expose that endpoint.
The Python Client
The prometheus_client library lets your Python service define and serve metrics with almost no boilerplate.
from prometheus_client import Counter, Histogram, start_http_serverA Counter Counts Up
A Counter only ever increases. It is perfect for totals like how many predictions your model has served since startup.
predictions = Counter("predictions_total", "Total predictions served")
predictions.inc()A Histogram Times Things
A Histogram records a distribution, ideal for request duration. It buckets values so you can later read p50 and p99 latency.
latency = Histogram("predict_seconds", "Prediction latency in seconds")
with latency.time():
model.predict(x)Labels Add Dimensions
Labels split one metric by category, like model version or status. Then you can compare error rates per version side by side.
errors = Counter("errors_total", "Errors", ["model"])
errors.labels(model="churn-v3").inc()Expose the /metrics Endpoint
Prometheus scrapes a path conventionally named /metrics. The client can serve it for you on its own port.
start_http_server(8000)
# metrics now live at http://localhost:8000/metricsWhat a Metric Looks Like
On that endpoint each metric is plain text: a name, optional labels, and a value. Prometheus reads this format on every scrape.
predictions_total 1423.0
predict_seconds_bucket{le="0.1"} 1390.0Configure the Scrape
Tell Prometheus where to look by adding your service as a target in its config. Then it pulls metrics on a fixed interval.
scrape_configs:
- job_name: model-api
static_configs:
- targets: ["model-api:8000"]Pick the Right Type
Rule of thumb: use a Counter for things that only grow, a Histogram for durations, and a Gauge for values that go up and down.
Quick Check
You want to track request latency so you can later read p99. Which Prometheus metric type fits?
Recap
You exposed a /metrics endpoint with prometheus_client, using Counters and Histograms plus labels. Prometheus now pulls live numbers from your model service. ✅
常见问题解答
「使用 Prometheus 暴露指标」课时是免费的吗?
是的 — 「使用 Prometheus 暴露指标」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 MLOps Academy 课程的其余内容,请升级到 CoddyKit PRO。 MLOps Academy 课程共包含 4 节课。
「使用 Prometheus 暴露指标」这节课中我会学到什么?
跟踪请求速率、错误和持续时间 你通过在浏览器中直接运行的动手代码来练习 MLOps Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 MLOps Academy 需要有经验吗?
无需任何先前经验。CoddyKit 上的 MLOps Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 2 节课,共 4 节。
「使用 Prometheus 暴露指标」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 MLOps Academy 课中编写并运行代码吗?
能。每节 MLOps Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- 用于预测的结构化日志
- 使用 Prometheus 暴露指标
- 构建 Grafana 仪表板
- 针对延迟和错误峰值发出警报