0Pricing
Pandas & NumPy Academy · 课时

检测缺失值

使用 isna()、notna() 和 isnull() 查找 Series 或 DataFrame 中 NaN 的位置,并统计每列的缺失值数量。

检测缺失值 是 CoddyKit 上的免费 Pandas & NumPy Academy 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Pandas & NumPy Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Pandas & NumPy Academy 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

What Are Missing Values in Pandas?

In Pandas, a missing value is represented as NaN (Not a Number) for numeric columns and None or pd.NaT for datetime columns. Missing data is common in real-world datasets because records may be incomplete, sensors may fail, or joins may produce unmatched rows. Detecting missing values is always the first step in any data cleaning workflow.

Pandas normalises None, float('nan'), and numpy.nan to the same internal NaN representation for numeric columns.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'name': ['Alice', None, 'Carol'],
    'age': [25, np.nan, 30],
    'salary': [50000.0, 60000.0, np.nan]
})
print(df)
#     name   age   salary
# 0  Alice  25.0  50000.0
# 1   None   NaN  60000.0
# 2  Carol  30.0      NaN

isna() and isnull()

isna() and isnull() are completely identical — both return a DataFrame or Series of the same shape filled with True wherever the value is missing and False elsewhere. Pandas provides both names purely for user preference. The result can be used directly as a boolean mask for filtering or further computation.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'A': [1, np.nan, 3],
    'B': [np.nan, 2, np.nan]
})

print(df.isna())
#        A      B
# 0  False   True
# 1   True  False
# 2  False   True

# Both are identical
print((df.isna() == df.isnull()).all().all())  # True

notna() to Find Non-Missing Values

notna() (also aliased as notnull()) is the inverse of isna() — it returns True where values are present and False where they are missing. This is useful when you want to filter to rows that have a value in a critical column, such as requiring that a primary key or target variable is not NaN.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'id': [1, 2, 3],
    'email': ['a@x.com', None, 'c@x.com']
})

# Keep only rows where email is present
with_email = df[df['email'].notna()]
print(with_email)
#    id    email
# 0   1  a@x.com
# 2   3  c@x.com

Counting Missing Values per Column

Calling .isna().sum() on a DataFrame sums the True values (which equal 1) column-by-column, giving you the count of missing values per column. This is the single most useful first step in understanding the quality of a new dataset — it tells you which columns need attention.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'name': ['Alice', None, 'Carol', None],
    'age': [25, np.nan, 30, 22],
    'salary': [50000, 60000, np.nan, np.nan]
})

missing_counts = df.isna().sum()
print(missing_counts)
# name      2
# age       1
# salary    2
# dtype: int64

Missing Percentage per Column

An absolute count of NaN values is less useful than the percentage of missing values because it scales with dataset size. Dividing isna().sum() by the total row count (or calling isna().mean()) gives the fraction missing, which you can multiply by 100 for a percentage. Columns with more than 30-50% missing often require a decision about whether to keep them at all.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'A': [1, np.nan, 3, np.nan, 5],
    'B': [np.nan, 2, np.nan, 4, 5],
    'C': [1, 2, 3, 4, 5]
})

missing_pct = (df.isna().mean() * 100).round(1)
print(missing_pct)
# A    40.0
# B    40.0
# C     0.0
# dtype: float64

Missing Summary Table

A common EDA pattern is to build a missing value summary table that shows count, percentage, and dtype for each column in one view. This gives a complete picture of data quality before making any cleaning decisions. You can sort it to surface the most problematic columns first.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'col_a': [1, np.nan, 3],
    'col_b': [np.nan, np.nan, 3],
    'col_c': [1, 2, 3]
})

summary = pd.DataFrame({
    'missing_count': df.isna().sum(),
    'missing_pct': (df.isna().mean() * 100).round(1),
    'dtype': df.dtypes
}).sort_values('missing_pct', ascending=False)
print(summary)
#         missing_count  missing_pct   dtype
# col_b               2         66.7  float64
# col_a               1         33.3  float64
# col_c               0          0.0  float64

Row-Level Missing Value Count

You can also count missing values per row by calling isna().sum(axis=1). This helps identify records that are mostly empty (e.g., incomplete survey responses) which you might want to flag or remove as a unit. A row with many missing values is fundamentally different from scattered column-level missingness.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'a': [1, np.nan, np.nan],
    'b': [np.nan, 2, np.nan],
    'c': [3, 4, np.nan]
})

# Count NaN per row
df['missing_count'] = df.isna().sum(axis=1)
print(df)
#      a    b    c  missing_count
# 0  1.0  NaN  3.0              1
# 1  NaN  2.0  4.0              1
# 2  NaN  NaN  NaN              3

Filtering Rows with Any or All Missing

Use df[df.isna().any(axis=1)] to find rows that have at least one NaN value, or df[df.isna().all(axis=1)] to find rows where every value is NaN. These filters help you isolate problem records for inspection before deciding how to handle them.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'x': [1, np.nan, np.nan],
    'y': [2, 3, np.nan],
    'z': [4, 5, np.nan]
})

# Rows with at least one NaN
has_any_nan = df[df.isna().any(axis=1)]
print('Any NaN:\n', has_any_nan)

# Rows where all values are NaN
all_nan = df[df.isna().all(axis=1)]
print('All NaN:\n', all_nan)
#      x    y    z
# 2  NaN  NaN  NaN

Checking a Specific Column for NaN

For a quick sanity check on a single column, call df['col'].isna().sum() or use df['col'].isna().any() to get a single boolean (True if any NaN exists). These one-liners are useful inside data validation checks or logging statements in a pipeline.

import pandas as pd
import numpy as np

df = pd.DataFrame({'price': [10.0, np.nan, 30.0, np.nan, 50.0]})

print('NaN count in price:', df['price'].isna().sum())  # 2
print('Any NaN in price?', df['price'].isna().any())    # True
print('All present?', df['price'].notna().all())         # False

Visualising Missing Values with a Heatmap

For datasets with many columns, a missing value heatmap is more informative than a table of numbers. You can create one easily with Seaborn: the darker a cell, the more missing data in that column-row combination. A popular third-party library called missingno provides dedicated missing-value visualisations.

import pandas as pd
import numpy as np
import seaborn as sns
import matplotlib.pyplot as plt

np.random.seed(0)
df = pd.DataFrame(
    np.where(np.random.rand(20, 5) > 0.7, np.nan, np.random.randn(20, 5)),
    columns=['A', 'B', 'C', 'D', 'E']
)

# Heatmap of missing values
sns.heatmap(df.isna(), cbar=False, yticklabels=False)
plt.title('Missing Value Pattern')
plt.tight_layout()
plt.savefig('missing_heatmap.png')
print('Saved missing_heatmap.png')

info() for a Quick Missing Check

df.info() prints a concise summary that includes the non-null count for every column. This is the fastest way to spot columns with missing values in a new dataset: any column whose non-null count is less than the total row count has NaN values. It also shows dtype and memory usage.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'id': [1, 2, 3, 4, 5],
    'age': [25, np.nan, 30, np.nan, 22],
    'salary': [50000, 60000, np.nan, 70000, 80000]
})

df.info()
# <class 'pandas.core.frame.DataFrame'>
# RangeIndex: 5 entries, 0 to 4
# Data columns (total 3 columns):
#  #   Column  Non-Null Count  Dtype
# ---  ------  --------------  -----
#  0   id      5 non-null      int64
#  1   age     3 non-null      float64
#  2   salary  4 non-null      float64

Quick Check

Test your understanding of detecting missing values in Pandas.

Lesson Recap

In this lesson you learned: isna() and isnull() are identical and return boolean masks of missing positions, isna().sum() counts NaN per column, and isna().mean()*100 gives the missing percentage. Use df.info() for a fast overview and isna().any(axis=1) to find rows with at least one NaN. Next up we tackle dropping missing values with dropna().

常见问题解答

「检测缺失值」课时是免费的吗?

是的 — 「检测缺失值」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Pandas & NumPy Academy 课程的其余内容,请升级到 CoddyKit PRO。 Pandas & NumPy Academy 课程共包含 4 节课。

「检测缺失值」这节课中我会学到什么?

使用 isna()、notna() 和 isnull() 查找 Series 或 DataFrame 中 NaN 的位置,并统计每列的缺失值数量。 你通过在浏览器中直接运行的动手代码来练习 Pandas & NumPy Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Pandas & NumPy Academy 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Pandas & NumPy Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。

「检测缺失值」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Pandas & NumPy Academy 课中编写并运行代码吗?

能。每节 Pandas & NumPy Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 检测缺失值
  2. 删除缺失值
  3. 填充缺失值
  4. 插值与高级缺失值填补
← 返回 Pandas & NumPy Academy