0Pricing
Data Science Academy · Lektion

K-Fold-Kreuzvalidierung

Scores über mehrere Folds mitteln

K-Fold-Kreuzvalidierung ist eine kostenlose Data Science Academy-Lektion auf CoddyKit. Dies ist Lektion 3 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des Data Science Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der Data Science Academy-Kurs umfasst insgesamt 4 Lektionen.

Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.

One Split, One Worry

A single train/test split gives just one score. If that split was lucky or unlucky, your estimate could be misleading.

The Big Idea

Cross-validation reuses your data many times, testing on a different slice each round, then averages the scores for a steadier estimate.

Cut Into Folds

You chop the data into k equal parts called folds. A common choice is five folds, giving five separate test rounds.

Rotate the Test Fold

Each round, one fold becomes the test set and the other k minus one folds train the model. Then the test fold rotates to the next one.

Everyone Gets a Turn

Over k rounds, every row is tested on exactly once and trained on the rest of the time. No data is wasted. 🔄

Average the Scores

You end with k scores, one per fold. Their average is your headline estimate, and their spread shows how stable the model is.

The Quick Way

scikit-learn does the looping for you. The helper cross_val_score returns one score per fold in a single line.

from sklearn.model_selection import cross_val_score
scores = cross_val_score(model, X, y, cv=5)

Read the Result

The returned array holds each fold's score. Take its mean for the headline and its standard deviation to gauge reliability.

print(scores.mean(), scores.std())

Choosing k

Five or ten folds are typical. More folds train on more data per round but cost more compute, so it is a speed-versus-stability trade.

Keep Classes Balanced

For classification, use StratifiedKFold so each fold mirrors the overall class balance. scikit-learn applies it automatically for classifiers.

Hold Out a Final Test

Cross-validation guides model choice during development. Still keep one untouched test set aside for a single honest score at the very end.

Quick Check

In 5-fold cross-validation, how often is each row tested?

Recap

Split into k folds, rotate the test fold, and average the scores for a robust estimate. Use cross_val_score, then a final hold-out. 🔄

Häufig gestellte Fragen

Ist die Lektion „K-Fold-Kreuzvalidierung“ kostenlos?

Ja — der vollständige Text von „K-Fold-Kreuzvalidierung“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des Data Science Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der Data Science Academy-Kurs umfasst insgesamt 4 Lektionen.

Was lerne ich in „K-Fold-Kreuzvalidierung“?

Scores über mehrere Folds mitteln Du übst Data Science Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.

Brauche ich Erfahrung, um Data Science Academy zu starten?

Keine Vorkenntnisse erforderlich. Data Science Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 3 von 4.

Wie lange dauert die Lektion „K-Fold-Kreuzvalidierung“?

Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.

Kann ich in dieser Data Science Academy-Lektion Code schreiben und ausführen?

Ja. Jede Data Science Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.

Alle Lektionen in diesem Kurs

  1. Warum Sie einen Testsatz zurückhalten
  2. train_test_split richtig einsetzen
  3. K-Fold-Kreuzvalidierung
  4. Datenlecks verhindern, bevor sie entstehen
← Zurück zu Data Science Academy