0Pricing
Data Science Academy · Lektion

Mit Mittelwert, Median oder Modus imputieren

Sinnvolle Ersetzungen je nach Spaltentyp

Mit Mittelwert, Median oder Modus imputieren ist eine kostenlose Data Science Academy-Lektion auf CoddyKit. Dies ist Lektion 3 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des Data Science Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der Data Science Academy-Kurs umfasst insgesamt 4 Lektionen.

Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.

Filling With a Statistic

Instead of inventing random numbers, you fill gaps with a summary value from the column itself. This careful approach is called imputation. 🧮

The Mean Fill

For numbers, the simplest choice is the mean, the column average. It keeps the overall total roughly intact when gaps are few.

df['age'].fillna(df['age'].mean())

When the Mean Misleads

The mean is fragile: a few huge values drag it far from the center. On skewed data, mean imputation can pull every gap toward an unrealistic number.

The Median Fill

The median is the middle value, and it shrugs off extreme outliers. For skewed numeric columns it is usually the safer fill.

df['income'].fillna(df['income'].median())

Mean or Median?

A quick rule: reach for the median when a column has outliers or a long tail, and the mean only when values are fairly symmetric.

The Mode for Categories

Text and category columns have no average. For them you fill with the mode, the most frequent value, since that is the most likely fit.

top = df['city'].mode()[0]
df['city'].fillna(top)

Why mode Returns a Series

A column can tie for most common, so mode() returns a Series of all winners. Pick the first with index 0 when you need one value.

df['city'].mode()[0]

Fill Per Column Type

The best practice is to impute each column by its type: median for skewed numbers, mean for symmetric ones, and mode for categories.

Forward and Back Fill

For ordered data like time series, carry the last known value forward with ffill, or pull the next one backward with bfill.

df['temp'].ffill()

The Hidden Cost

Every imputation shrinks the column's natural spread, because filled values cluster at one point. Note this, since it nudges your variance downward.

Flag What You Filled

A pro habit: add a boolean column marking which rows were imputed. That flag lets later analysis know which values were real and which were guessed.

df['age_filled'] = df['age'].isna()

Quick Check

An income column is heavily skewed by a few millionaires. Which imputation fits best?

Recap: Smart Fills

You now impute with mean, median, or mode by column type, use ffill for ordered data, and flag filled rows so nothing gets hidden.

Häufig gestellte Fragen

Ist die Lektion „Mit Mittelwert, Median oder Modus imputieren“ kostenlos?

Ja — der vollständige Text von „Mit Mittelwert, Median oder Modus imputieren“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des Data Science Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der Data Science Academy-Kurs umfasst insgesamt 4 Lektionen.

Was lerne ich in „Mit Mittelwert, Median oder Modus imputieren“?

Sinnvolle Ersetzungen je nach Spaltentyp Du übst Data Science Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.

Brauche ich Erfahrung, um Data Science Academy zu starten?

Keine Vorkenntnisse erforderlich. Data Science Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 3 von 4.

Wie lange dauert die Lektion „Mit Mittelwert, Median oder Modus imputieren“?

Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.

Kann ich in dieser Data Science Academy-Lektion Code schreiben und ausführen?

Ja. Jede Data Science Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.

Alle Lektionen in diesem Kurs

  1. Die verborgenen NaNs in Ihrer Tabelle finden
  2. Löschen oder auffüllen: mit Bedacht wählen
  3. Mit Mittelwert, Median oder Modus imputieren
  4. dtypes und doppelte Zeilen korrigieren
← Zurück zu Data Science Academy