0Pricing
Pandas & NumPy Academy · Leçon

Supprimer les valeurs manquantes

Supprimez les lignes ou les colonnes contenant NaN avec dropna(), en contrôlant le seuil et le sous-ensemble de colonnes pris en compte.

Supprimer les valeurs manquantes est une leçon Pandas & NumPy Academy gratuite sur CoddyKit. Ceci est la leçon 2 sur 4. Tu peux lire la leçon complète ci-dessous gratuitement — puis la pratiquer en direct dans le navigateur avec un éditeur de code intégré et un tuteur IA 24/7. Elle fait partie du parcours d'apprentissage Pandas & NumPy Academy, et ta progression se synchronise sur le web et l'application CoddyKit. Le cours Pandas & NumPy Academy comprend 4 leçons au total.

Certaines parties de cette leçon n'ont pas encore été traduites et s'affichent en anglais.

When to Drop Missing Values?

Dropping rows with missing values is the simplest imputation strategy, but it is only valid when the data is Missing Completely At Random (MCAR) — meaning the probability of a value being missing has nothing to do with the missing value itself or any other variable. If missing data is systematic (e.g., low-income respondents skip the salary field), dropping it introduces bias. Always investigate the missingness pattern before deciding to drop.

import pandas as pd
import numpy as np

# Example: randomly missing salary data (MCAR-like)
df = pd.DataFrame({
    'name': ['Alice', 'Bob', 'Carol', 'Dave'],
    'salary': [50000, np.nan, 70000, np.nan]
})
print('Before drop:', df.shape)  # (4, 2)
print(df)

dropna() — Basic Usage

DataFrame.dropna() removes any row that contains at least one NaN value by default. It returns a new DataFrame; the original is unchanged unless you pass inplace=True. For small to medium datasets this default behaviour is often acceptable as a quick data cleaning first pass.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'a': [1, np.nan, 3, np.nan],
    'b': [10, 20, np.nan, 40],
    'c': [100, 200, 300, 400]
})

cleaned = df.dropna()
print(cleaned)
#      a     b    c
# 0  1.0  10.0  100

print('Original shape:', df.shape)    # (4, 3)
print('Cleaned shape:', cleaned.shape) # (1, 3)

how='all' — Drop Only All-NaN Rows

Passing how='all' tells dropna to remove a row only if every single value in that row is NaN. This is much less aggressive than the default how='any'. Use how='all' when your dataset has sparse rows — rows that have some data are worth keeping, while entirely empty rows are clearly junk records.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'a': [1, np.nan, np.nan],
    'b': [2, np.nan, np.nan],
    'c': [3, 4, np.nan]
})

# Row 2 has ALL NaN — dropped
# Row 1 has partial NaN — kept with how='all'
cleaned = df.dropna(how='all')
print(cleaned)
#      a    b    c
# 0  1.0  2.0  3.0
# 1  NaN  NaN  4.0

subset= — Check Only Specific Columns

The subset parameter limits which columns are checked for NaN when deciding whether to drop a row. This is extremely useful when only certain columns are critical — for example, a row should be dropped if the user_id or target column is missing, but NaN in optional feature columns is acceptable.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'user_id': [1, np.nan, 3],
    'score': [88, 95, np.nan],
    'notes': ['ok', 'good', np.nan]
})

# Drop row only if user_id is missing — score and notes NaN are ok
cleaned = df.dropna(subset=['user_id'])
print(cleaned)
#    user_id  score notes
# 0      1.0   88.0    ok
# 2      3.0    NaN   NaN

thresh= — Minimum Non-Null Requirement

The thresh parameter keeps a row only if it has at least thresh non-NaN values. This is more nuanced than how='any' or how='all': you can say 'keep a row if at least 3 out of 5 columns have data'. This is useful for datasets where some sparsity is expected but completely empty rows should be removed.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'a': [1, np.nan, np.nan],
    'b': [2, 3, np.nan],
    'c': [4, np.nan, np.nan],
    'd': [5, 6, np.nan]
})

# Keep rows with at least 3 non-null values
cleaned = df.dropna(thresh=3)
print(cleaned)
#      a    b    c    d
# 0  1.0  2.0  4.0  5.0
# 1  NaN  3.0  NaN  6.0

Dropping Columns Instead of Rows

By default, dropna() removes rows (axis=0). Pass axis=1 (or axis='columns') to drop columns that contain any NaN instead. This is appropriate when a column is mostly empty and provides little signal — keeping it would just add noise to a model or summary table.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'id': [1, 2, 3],
    'name': ['A', 'B', 'C'],
    'temp': [np.nan, np.nan, np.nan],  # completely empty column
    'score': [80, 90, 85]
})

# Drop columns that have any NaN
cleaned = df.dropna(axis=1)
print(cleaned)
#    id name  score
# 0   1    A     80
# 1   2    B     90
# 2   3    C     85

Dropping Columns by Missing Threshold

A powerful pattern is to drop columns that exceed a certain missing percentage. Compute the fraction missing per column, identify columns above your threshold (e.g., 50%), and drop them with df.drop(columns=cols_to_drop). This is more targeted than dropna(axis=1) which drops any column with even one NaN.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'a': [1, 2, np.nan, 4, 5],
    'b': [np.nan, np.nan, np.nan, np.nan, 5],  # 80% missing
    'c': [1, np.nan, 3, 4, 5]                   # 20% missing
})

threshold = 0.5
high_missing = df.columns[df.isna().mean() > threshold]
print('Dropping:', high_missing.tolist())  # ['b']

cleaned = df.drop(columns=high_missing)
print(cleaned)

Preserving the Index After dropna()

After calling dropna(), the original row indices are preserved — so if rows 1 and 3 were dropped, the remaining DataFrame has indices 0, 2, 4. This is often desirable (you can trace back to original positions), but sometimes you want a clean sequential index starting from 0. Call .reset_index(drop=True) after dropping to renumber rows.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'val': [1, np.nan, 3, np.nan, 5]
})

cleaned = df.dropna()
print('With original index:')
print(cleaned)  # indices 0, 2, 4

cleaned_reset = cleaned.reset_index(drop=True)
print('With reset index:')
print(cleaned_reset)  # indices 0, 1, 2

dropna() in a Pipeline

Since dropna() returns a DataFrame, it integrates naturally into a method chain. Chaining dropna() between load and analysis steps is a clean pattern that keeps the pipeline readable without temporary variables. You can also chain it with query(), assign(), and groupby().

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'region': ['East', 'West', None, 'East'],
    'revenue': [100, np.nan, 300, 400]
})

result = (
    df
    .dropna(subset=['region', 'revenue'])
    .groupby('region')['revenue'].sum()
)
print(result)
# region
# East    500.0
# dtype: float64

When NOT to Drop: Prefer Filling

Dropping rows loses data. For columns with fewer than 5-10% missing values, filling (imputing) is usually better than dropping. Also, if missing values are correlated with the target variable (Missing Not At Random, MNAR), dropping them introduces bias. As a rule: only drop when missingness is truly random, the dataset is large enough that lost rows don't matter, and the column or row provides no recoverable signal.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'user': ['Alice', 'Bob', 'Carol', 'Dave'],
    'age': [25, np.nan, 30, np.nan]
})

missing_pct = df['age'].isna().mean()
print(f'Missing age: {missing_pct:.0%}')  # 50%
# 50% missing is high — consider imputing instead of dropping
# df['age'].fillna(df['age'].median(), inplace=True)

Practical Cleaning Workflow

A practical missing-value workflow combines multiple dropna strategies: first drop entirely empty rows, then drop columns that are more than 60% empty, then drop rows missing critical ID or target columns, and finally fill the remaining scattered NaN values. This layered approach preserves as much data as possible.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'id': [1, 2, 3, np.nan],
    'feature': [1.0, np.nan, 3.0, 4.0],
    'useless': [np.nan, np.nan, np.nan, np.nan]
})

clean = (
    df
    .dropna(how='all')              # remove all-NaN rows
    .drop(columns=df.columns[df.isna().mean() > 0.9])  # remove >90% empty cols
    .dropna(subset=['id'])           # must have an ID
)
print(clean)
#      id  feature
# 0   1.0      1.0
# 1   2.0      NaN
# 2   3.0      3.0

Quick Check

Test your understanding of dropping missing values with dropna().

Lesson Recap

In this lesson you learned: dropna() removes rows with NaN by default, how='all' only drops fully-empty rows, subset= restricts checking to specific columns, and thresh= keeps rows with a minimum number of non-null values. Use axis=1 to drop columns instead of rows. Next up we fill missing values instead of dropping them using fillna().

Questions Fréquemment Posées

La leçon « Supprimer les valeurs manquantes » est-elle gratuite ?

Oui — le texte complet de « Supprimer les valeurs manquantes » est gratuit à lire ici sur le web. Pour la pratiquer de manière interactive (un éditeur de code intégré et un tuteur IA 24/7) et déverrouiller le reste du cours Pandas & NumPy Academy, passe à CoddyKit PRO. Le cours Pandas & NumPy Academy comprend 4 leçons au total.

Qu'est-ce que j'apprendrai dans « Supprimer les valeurs manquantes » ?

Supprimez les lignes ou les colonnes contenant NaN avec dropna(), en contrôlant le seuil et le sous-ensemble de colonnes pris en compte. Tu pratiques Pandas & NumPy Academy avec du code pratique que tu exécutes directement dans le navigateur, et un tuteur IA 24/7 répond à tes questions au fur et à mesure que tu avances dans la leçon.

Dois-je avoir de l'expérience pour commencer Pandas & NumPy Academy ?

Aucune expérience préalable n'est requise. Pandas & NumPy Academy sur CoddyKit est structuré pour les débutants jusqu'aux apprenants avancés, donc tu peux commencer ici ou depuis le début et avancer à ton rythme. Ceci est la leçon 2 sur 4.

Combien de temps prend la leçon « Supprimer les valeurs manquantes » ?

La plupart des leçons CoddyKit prennent environ 5–10 minutes. Chacune est courte et interactive, tu progresses régulièrement et tu repiques exactement où tu t'es arrêté sur le web et l'app.

Peux-tu écrire et exécuter du code dans cette leçon Pandas & NumPy Academy ?

Oui. Chaque leçon Pandas & NumPy Academy inclut un éditeur de code intégré, tu écris et exécutes du vrai code directement dans ton navigateur et tu reçois des retours IA instantanés — aucune configuration locale requise.

Toutes les leçons de ce cours

  1. Détecter les valeurs manquantes
  2. Supprimer les valeurs manquantes
  3. Remplacer les valeurs manquantes
  4. Interpolation et imputation avancée
← Retour à Pandas & NumPy Academy