0Pricing
MLOps Academy · 강의

데이터에서 모델까지 계보 추적하기

배포된 모델을 정확한 데이터와 코드까지 연결하세요.

데이터에서 모델까지 계보 추적하기은(는) CoddyKit의 무료 MLOps Academy 강의입니다. 이것은 4개 중 1번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 MLOps Academy 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. MLOps Academy 강의에는 총 4개의 강의가 포함되어 있습니다.

이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.

What Lineage Means

Lineage is the documented chain that links a deployed model back to the exact data, code, and config that produced it. 🔗

Why You Need It

When a prediction is questioned, lineage lets you answer one hard question: which data and which code version created this model?

The Three Inputs to Track

Every trained model has three parents worth tracking: the dataset version, the training code commit, and the run parameters.

Pin the Code Commit

Log the exact Git commit hash with each run so you always know which code trained the model.

import subprocess
commit = subprocess.check_output(["git", "rev-parse", "HEAD"]).decode().strip()
mlflow.set_tag("git_commit", commit)

Pin the Data Version

Record the data version too. With DVC, the data hash lives in Git, so the commit already points to one exact dataset state.

MLflow Stores the Link

MLflow saves params, metrics, and tags per run, so a logged model already carries pointers to how it was built. 🧾

mlflow.log_param("data_path", "data/train.csv")
mlflow.log_metric("f1", 0.91)
mlflow.sklearn.log_model(model, "model")

Tag the Run Generously

Tags are free metadata. Attach the dataset hash, environment name, and author so the run tells its own story later.

mlflow.set_tag("dataset_hash", "a1b2c3")
mlflow.set_tag("author", "team-fraud")

Lineage Is a Graph

Picture lineage as a graph: data nodes flow into a training run, which flows into a model, which flows into a deployment.

Trace Forward and Backward

Good lineage works both ways. Backward answers what built this model; forward answers which models a dataset affected.

Read Lineage From a Run

You can fetch a past run by id and read its tags to reconstruct exactly which data and commit produced that model.

run = mlflow.get_run(run_id)
print(run.data.tags["git_commit"])
print(run.data.tags["dataset_hash"])

Tools That Do This

MLflow, DVC, and dedicated tools like OpenLineage capture these links automatically so you do not stitch lineage together by hand.

Quick Check

Test your grasp of what lineage actually links together.

Recap

You learned that lineage ties a model to its data, code, and params, captured as tags on a run so any model is fully traceable. ✅

자주 묻는 질문

“데이터에서 모델까지 계보 추적하기” 강의는 무료인가요?

네 — “데이터에서 모델까지 계보 추적하기” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 MLOps Academy 강의 전체를 잠금 해제할 수 있습니다. MLOps Academy 강의에는 총 4개의 강의가 포함되어 있습니다.

“데이터에서 모델까지 계보 추적하기”에서 뭘 배우나요?

배포된 모델을 정확한 데이터와 코드까지 연결하세요. 브라우저에서 직접 실행하는 실습 코드로 MLOps Academy을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.

MLOps Academy을(를) 시작하는 데 경험이 필요한가요?

사전 경험은 필요하지 않습니다. CoddyKit의 MLOps Academy은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 1번째 강의입니다.

“데이터에서 모델까지 계보 추적하기” 강의는 얼마나 걸리나요?

대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.

이 MLOps Academy 강의에서 코드를 작성하고 실행할 수 있나요?

네. 모든 MLOps Academy 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.

이 강의의 모든 강의

  1. 데이터에서 모델까지 계보 추적하기
  2. 모델 카드 작성하기
  3. 감사 추적과 재현성
  4. 접근 제어와 규정 준수
← MLOps Academy(으)로 돌아가기