0Pricing
Pandas & NumPy Academy · Урок

Удаление пропущенных значений

Удаляйте строки или столбцы, содержащие NaN, с помощью dropna(), задавая порог и набор учитываемых столбцов.

«Удаление пропущенных значений» — бесплатный урок Pandas & NumPy Academy на CoddyKit. Это урок 2 из 4. Ты можешь прочитать весь урок бесплатно ниже — а потом практиковать его прямо в браузере с встроенным редактором кода и ИИ-репетитором 24/7. Это часть пути обучения Pandas & NumPy Academy, и твой прогресс синхронизируется между веб-версией и приложением CoddyKit. Курс Pandas & NumPy Academy содержит 4 уроков всего.

Части этого урока еще не переведены и отображаются на английском.

When to Drop Missing Values?

Dropping rows with missing values is the simplest imputation strategy, but it is only valid when the data is Missing Completely At Random (MCAR) — meaning the probability of a value being missing has nothing to do with the missing value itself or any other variable. If missing data is systematic (e.g., low-income respondents skip the salary field), dropping it introduces bias. Always investigate the missingness pattern before deciding to drop.

import pandas as pd
import numpy as np

# Example: randomly missing salary data (MCAR-like)
df = pd.DataFrame({
    'name': ['Alice', 'Bob', 'Carol', 'Dave'],
    'salary': [50000, np.nan, 70000, np.nan]
})
print('Before drop:', df.shape)  # (4, 2)
print(df)

dropna() — Basic Usage

DataFrame.dropna() removes any row that contains at least one NaN value by default. It returns a new DataFrame; the original is unchanged unless you pass inplace=True. For small to medium datasets this default behaviour is often acceptable as a quick data cleaning first pass.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'a': [1, np.nan, 3, np.nan],
    'b': [10, 20, np.nan, 40],
    'c': [100, 200, 300, 400]
})

cleaned = df.dropna()
print(cleaned)
#      a     b    c
# 0  1.0  10.0  100

print('Original shape:', df.shape)    # (4, 3)
print('Cleaned shape:', cleaned.shape) # (1, 3)

how='all' — Drop Only All-NaN Rows

Passing how='all' tells dropna to remove a row only if every single value in that row is NaN. This is much less aggressive than the default how='any'. Use how='all' when your dataset has sparse rows — rows that have some data are worth keeping, while entirely empty rows are clearly junk records.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'a': [1, np.nan, np.nan],
    'b': [2, np.nan, np.nan],
    'c': [3, 4, np.nan]
})

# Row 2 has ALL NaN — dropped
# Row 1 has partial NaN — kept with how='all'
cleaned = df.dropna(how='all')
print(cleaned)
#      a    b    c
# 0  1.0  2.0  3.0
# 1  NaN  NaN  4.0

subset= — Check Only Specific Columns

The subset parameter limits which columns are checked for NaN when deciding whether to drop a row. This is extremely useful when only certain columns are critical — for example, a row should be dropped if the user_id or target column is missing, but NaN in optional feature columns is acceptable.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'user_id': [1, np.nan, 3],
    'score': [88, 95, np.nan],
    'notes': ['ok', 'good', np.nan]
})

# Drop row only if user_id is missing — score and notes NaN are ok
cleaned = df.dropna(subset=['user_id'])
print(cleaned)
#    user_id  score notes
# 0      1.0   88.0    ok
# 2      3.0    NaN   NaN

thresh= — Minimum Non-Null Requirement

The thresh parameter keeps a row only if it has at least thresh non-NaN values. This is more nuanced than how='any' or how='all': you can say 'keep a row if at least 3 out of 5 columns have data'. This is useful for datasets where some sparsity is expected but completely empty rows should be removed.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'a': [1, np.nan, np.nan],
    'b': [2, 3, np.nan],
    'c': [4, np.nan, np.nan],
    'd': [5, 6, np.nan]
})

# Keep rows with at least 3 non-null values
cleaned = df.dropna(thresh=3)
print(cleaned)
#      a    b    c    d
# 0  1.0  2.0  4.0  5.0
# 1  NaN  3.0  NaN  6.0

Dropping Columns Instead of Rows

By default, dropna() removes rows (axis=0). Pass axis=1 (or axis='columns') to drop columns that contain any NaN instead. This is appropriate when a column is mostly empty and provides little signal — keeping it would just add noise to a model or summary table.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'id': [1, 2, 3],
    'name': ['A', 'B', 'C'],
    'temp': [np.nan, np.nan, np.nan],  # completely empty column
    'score': [80, 90, 85]
})

# Drop columns that have any NaN
cleaned = df.dropna(axis=1)
print(cleaned)
#    id name  score
# 0   1    A     80
# 1   2    B     90
# 2   3    C     85

Dropping Columns by Missing Threshold

A powerful pattern is to drop columns that exceed a certain missing percentage. Compute the fraction missing per column, identify columns above your threshold (e.g., 50%), and drop them with df.drop(columns=cols_to_drop). This is more targeted than dropna(axis=1) which drops any column with even one NaN.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'a': [1, 2, np.nan, 4, 5],
    'b': [np.nan, np.nan, np.nan, np.nan, 5],  # 80% missing
    'c': [1, np.nan, 3, 4, 5]                   # 20% missing
})

threshold = 0.5
high_missing = df.columns[df.isna().mean() > threshold]
print('Dropping:', high_missing.tolist())  # ['b']

cleaned = df.drop(columns=high_missing)
print(cleaned)

Preserving the Index After dropna()

After calling dropna(), the original row indices are preserved — so if rows 1 and 3 were dropped, the remaining DataFrame has indices 0, 2, 4. This is often desirable (you can trace back to original positions), but sometimes you want a clean sequential index starting from 0. Call .reset_index(drop=True) after dropping to renumber rows.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'val': [1, np.nan, 3, np.nan, 5]
})

cleaned = df.dropna()
print('With original index:')
print(cleaned)  # indices 0, 2, 4

cleaned_reset = cleaned.reset_index(drop=True)
print('With reset index:')
print(cleaned_reset)  # indices 0, 1, 2

dropna() in a Pipeline

Since dropna() returns a DataFrame, it integrates naturally into a method chain. Chaining dropna() between load and analysis steps is a clean pattern that keeps the pipeline readable without temporary variables. You can also chain it with query(), assign(), and groupby().

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'region': ['East', 'West', None, 'East'],
    'revenue': [100, np.nan, 300, 400]
})

result = (
    df
    .dropna(subset=['region', 'revenue'])
    .groupby('region')['revenue'].sum()
)
print(result)
# region
# East    500.0
# dtype: float64

When NOT to Drop: Prefer Filling

Dropping rows loses data. For columns with fewer than 5-10% missing values, filling (imputing) is usually better than dropping. Also, if missing values are correlated with the target variable (Missing Not At Random, MNAR), dropping them introduces bias. As a rule: only drop when missingness is truly random, the dataset is large enough that lost rows don't matter, and the column or row provides no recoverable signal.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'user': ['Alice', 'Bob', 'Carol', 'Dave'],
    'age': [25, np.nan, 30, np.nan]
})

missing_pct = df['age'].isna().mean()
print(f'Missing age: {missing_pct:.0%}')  # 50%
# 50% missing is high — consider imputing instead of dropping
# df['age'].fillna(df['age'].median(), inplace=True)

Practical Cleaning Workflow

A practical missing-value workflow combines multiple dropna strategies: first drop entirely empty rows, then drop columns that are more than 60% empty, then drop rows missing critical ID or target columns, and finally fill the remaining scattered NaN values. This layered approach preserves as much data as possible.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'id': [1, 2, 3, np.nan],
    'feature': [1.0, np.nan, 3.0, 4.0],
    'useless': [np.nan, np.nan, np.nan, np.nan]
})

clean = (
    df
    .dropna(how='all')              # remove all-NaN rows
    .drop(columns=df.columns[df.isna().mean() > 0.9])  # remove >90% empty cols
    .dropna(subset=['id'])           # must have an ID
)
print(clean)
#      id  feature
# 0   1.0      1.0
# 1   2.0      NaN
# 2   3.0      3.0

Quick Check

Test your understanding of dropping missing values with dropna().

Lesson Recap

In this lesson you learned: dropna() removes rows with NaN by default, how='all' only drops fully-empty rows, subset= restricts checking to specific columns, and thresh= keeps rows with a minimum number of non-null values. Use axis=1 to drop columns instead of rows. Next up we fill missing values instead of dropping them using fillna().

Часто задаваемые вопросы

Урок «Удаление пропущенных значений» бесплатный?

Да — полный текст урока «Удаление пропущенных значений» бесплатно доступен здесь в веб-версии. Чтобы практиковать его интерактивно (встроенный редактор кода и ИИ-репетитор 24/7) и разблокировать остальной курс Pandas & NumPy Academy, подпишись на CoddyKit PRO. Курс Pandas & NumPy Academy содержит 4 уроков всего.

Чему я научусь в уроке «Удаление пропущенных значений»?

Удаляйте строки или столбцы, содержащие NaN, с помощью dropna(), задавая порог и набор учитываемых столбцов. Ты практикуешь Pandas & NumPy Academy с помощью реального кода, который запускаешь прямо в браузере, и ИИ-репетитор 24/7 отвечает на твои вопросы во время урока.

Нужен ли мне опыт, чтобы начать Pandas & NumPy Academy?

Предыдущий опыт не требуется. Pandas & NumPy Academy на CoddyKit структурирован для всех уровней — от новичков до продвинутых, поэтому ты можешь начать отсюда или с самого начала и учиться в своем темпе. Это урок 2 из 4.

Сколько времени занимает урок «Удаление пропущенных значений»?

Большинство уроков CoddyKit занимают около 5–10 минут. Каждый из них компактный и интерактивный, поэтому ты постоянно делаешь прогресс и продолжаешь с того же места в веб-версии и приложении.

Можно ли писать и запускать код в этом уроке Pandas & NumPy Academy?

Да. Каждый урок Pandas & NumPy Academy включает встроенный редактор кода, поэтому ты пишешь и запускаешь реальный код прямо в браузере и получаешь моментальную обратную связь от AI — локальная установка не требуется.

Все уроки этого курса

  1. Поиск пропущенных значений
  2. Удаление пропущенных значений
  3. Заполнение пропущенных значений
  4. Интерполяция и продвинутое заполнение
← Назад к Pandas & NumPy Academy