Eliminar valores faltantes
Elimine las filas o columnas que contengan NaN con dropna(), controlando el umbral y el subconjunto de columnas considerados.
Eliminar valores faltantes es una lección gratuita de Pandas & NumPy Academy en CoddyKit. Esta es la lección 2 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Pandas & NumPy Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Pandas & NumPy Academy incluye 4 lecciones en total.
Partes de esta lección aún no han sido traducidas y se muestran en inglés.
When to Drop Missing Values?
Dropping rows with missing values is the simplest imputation strategy, but it is only valid when the data is Missing Completely At Random (MCAR) — meaning the probability of a value being missing has nothing to do with the missing value itself or any other variable. If missing data is systematic (e.g., low-income respondents skip the salary field), dropping it introduces bias. Always investigate the missingness pattern before deciding to drop.
import pandas as pd
import numpy as np
# Example: randomly missing salary data (MCAR-like)
df = pd.DataFrame({
'name': ['Alice', 'Bob', 'Carol', 'Dave'],
'salary': [50000, np.nan, 70000, np.nan]
})
print('Before drop:', df.shape) # (4, 2)
print(df)dropna() — Basic Usage
DataFrame.dropna() removes any row that contains at least one NaN value by default. It returns a new DataFrame; the original is unchanged unless you pass inplace=True. For small to medium datasets this default behaviour is often acceptable as a quick data cleaning first pass.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'a': [1, np.nan, 3, np.nan],
'b': [10, 20, np.nan, 40],
'c': [100, 200, 300, 400]
})
cleaned = df.dropna()
print(cleaned)
# a b c
# 0 1.0 10.0 100
print('Original shape:', df.shape) # (4, 3)
print('Cleaned shape:', cleaned.shape) # (1, 3)how='all' — Drop Only All-NaN Rows
Passing how='all' tells dropna to remove a row only if every single value in that row is NaN. This is much less aggressive than the default how='any'. Use how='all' when your dataset has sparse rows — rows that have some data are worth keeping, while entirely empty rows are clearly junk records.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'a': [1, np.nan, np.nan],
'b': [2, np.nan, np.nan],
'c': [3, 4, np.nan]
})
# Row 2 has ALL NaN — dropped
# Row 1 has partial NaN — kept with how='all'
cleaned = df.dropna(how='all')
print(cleaned)
# a b c
# 0 1.0 2.0 3.0
# 1 NaN NaN 4.0subset= — Check Only Specific Columns
The subset parameter limits which columns are checked for NaN when deciding whether to drop a row. This is extremely useful when only certain columns are critical — for example, a row should be dropped if the user_id or target column is missing, but NaN in optional feature columns is acceptable.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'user_id': [1, np.nan, 3],
'score': [88, 95, np.nan],
'notes': ['ok', 'good', np.nan]
})
# Drop row only if user_id is missing — score and notes NaN are ok
cleaned = df.dropna(subset=['user_id'])
print(cleaned)
# user_id score notes
# 0 1.0 88.0 ok
# 2 3.0 NaN NaNthresh= — Minimum Non-Null Requirement
The thresh parameter keeps a row only if it has at least thresh non-NaN values. This is more nuanced than how='any' or how='all': you can say 'keep a row if at least 3 out of 5 columns have data'. This is useful for datasets where some sparsity is expected but completely empty rows should be removed.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'a': [1, np.nan, np.nan],
'b': [2, 3, np.nan],
'c': [4, np.nan, np.nan],
'd': [5, 6, np.nan]
})
# Keep rows with at least 3 non-null values
cleaned = df.dropna(thresh=3)
print(cleaned)
# a b c d
# 0 1.0 2.0 4.0 5.0
# 1 NaN 3.0 NaN 6.0Dropping Columns Instead of Rows
By default, dropna() removes rows (axis=0). Pass axis=1 (or axis='columns') to drop columns that contain any NaN instead. This is appropriate when a column is mostly empty and provides little signal — keeping it would just add noise to a model or summary table.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'id': [1, 2, 3],
'name': ['A', 'B', 'C'],
'temp': [np.nan, np.nan, np.nan], # completely empty column
'score': [80, 90, 85]
})
# Drop columns that have any NaN
cleaned = df.dropna(axis=1)
print(cleaned)
# id name score
# 0 1 A 80
# 1 2 B 90
# 2 3 C 85Dropping Columns by Missing Threshold
A powerful pattern is to drop columns that exceed a certain missing percentage. Compute the fraction missing per column, identify columns above your threshold (e.g., 50%), and drop them with df.drop(columns=cols_to_drop). This is more targeted than dropna(axis=1) which drops any column with even one NaN.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'a': [1, 2, np.nan, 4, 5],
'b': [np.nan, np.nan, np.nan, np.nan, 5], # 80% missing
'c': [1, np.nan, 3, 4, 5] # 20% missing
})
threshold = 0.5
high_missing = df.columns[df.isna().mean() > threshold]
print('Dropping:', high_missing.tolist()) # ['b']
cleaned = df.drop(columns=high_missing)
print(cleaned)Preserving the Index After dropna()
After calling dropna(), the original row indices are preserved — so if rows 1 and 3 were dropped, the remaining DataFrame has indices 0, 2, 4. This is often desirable (you can trace back to original positions), but sometimes you want a clean sequential index starting from 0. Call .reset_index(drop=True) after dropping to renumber rows.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'val': [1, np.nan, 3, np.nan, 5]
})
cleaned = df.dropna()
print('With original index:')
print(cleaned) # indices 0, 2, 4
cleaned_reset = cleaned.reset_index(drop=True)
print('With reset index:')
print(cleaned_reset) # indices 0, 1, 2dropna() in a Pipeline
Since dropna() returns a DataFrame, it integrates naturally into a method chain. Chaining dropna() between load and analysis steps is a clean pattern that keeps the pipeline readable without temporary variables. You can also chain it with query(), assign(), and groupby().
import pandas as pd
import numpy as np
df = pd.DataFrame({
'region': ['East', 'West', None, 'East'],
'revenue': [100, np.nan, 300, 400]
})
result = (
df
.dropna(subset=['region', 'revenue'])
.groupby('region')['revenue'].sum()
)
print(result)
# region
# East 500.0
# dtype: float64When NOT to Drop: Prefer Filling
Dropping rows loses data. For columns with fewer than 5-10% missing values, filling (imputing) is usually better than dropping. Also, if missing values are correlated with the target variable (Missing Not At Random, MNAR), dropping them introduces bias. As a rule: only drop when missingness is truly random, the dataset is large enough that lost rows don't matter, and the column or row provides no recoverable signal.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'user': ['Alice', 'Bob', 'Carol', 'Dave'],
'age': [25, np.nan, 30, np.nan]
})
missing_pct = df['age'].isna().mean()
print(f'Missing age: {missing_pct:.0%}') # 50%
# 50% missing is high — consider imputing instead of dropping
# df['age'].fillna(df['age'].median(), inplace=True)Practical Cleaning Workflow
A practical missing-value workflow combines multiple dropna strategies: first drop entirely empty rows, then drop columns that are more than 60% empty, then drop rows missing critical ID or target columns, and finally fill the remaining scattered NaN values. This layered approach preserves as much data as possible.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'id': [1, 2, 3, np.nan],
'feature': [1.0, np.nan, 3.0, 4.0],
'useless': [np.nan, np.nan, np.nan, np.nan]
})
clean = (
df
.dropna(how='all') # remove all-NaN rows
.drop(columns=df.columns[df.isna().mean() > 0.9]) # remove >90% empty cols
.dropna(subset=['id']) # must have an ID
)
print(clean)
# id feature
# 0 1.0 1.0
# 1 2.0 NaN
# 2 3.0 3.0Quick Check
Test your understanding of dropping missing values with dropna().
Lesson Recap
In this lesson you learned: dropna() removes rows with NaN by default, how='all' only drops fully-empty rows, subset= restricts checking to specific columns, and thresh= keeps rows with a minimum number of non-null values. Use axis=1 to drop columns instead of rows. Next up we fill missing values instead of dropping them using fillna().
Preguntas frecuentes
¿La lección «Eliminar valores faltantes» es gratis?
Sí — el texto completo de «Eliminar valores faltantes» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Pandas & NumPy Academy, actualiza a CoddyKit PRO. El curso de Pandas & NumPy Academy incluye 4 lecciones en total.
¿Qué aprenderé en «Eliminar valores faltantes»?
Elimine las filas o columnas que contengan NaN con dropna(), controlando el umbral y el subconjunto de columnas considerados. Practicas Pandas & NumPy Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.
¿Necesito experiencia previa para empezar Pandas & NumPy Academy?
No se requiere experiencia previa. Pandas & NumPy Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 2 de 4.
¿Cuánto tiempo toma la lección «Eliminar valores faltantes»?
La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.
¿Puedo escribir y ejecutar código en esta lección de Pandas & NumPy Academy?
Sí. Cada lección de Pandas & NumPy Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.
Todas las lecciones de este curso
- Detectar valores faltantes
- Eliminar valores faltantes
- Rellenar valores faltantes
- Interpolación e imputación avanzada