欠損値の検出
isna()、notna()、isnull()を使って、SeriesまたはDataFrame内のNaNの位置を見つけ、列ごとの欠損値を数えます。
「欠損値の検出」はCoddyKit上の無料Pandas & NumPy Academyレッスンです。 これはレッスン1/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはPandas & NumPy Academy学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Pandas & NumPy Academyコースには全4レッスンが含まれています。
このレッスンの一部はまだ翻訳されておらず、英語で表示されています。
What Are Missing Values in Pandas?
In Pandas, a missing value is represented as NaN (Not a Number) for numeric columns and None or pd.NaT for datetime columns. Missing data is common in real-world datasets because records may be incomplete, sensors may fail, or joins may produce unmatched rows. Detecting missing values is always the first step in any data cleaning workflow.
Pandas normalises None, float('nan'), and numpy.nan to the same internal NaN representation for numeric columns.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'name': ['Alice', None, 'Carol'],
'age': [25, np.nan, 30],
'salary': [50000.0, 60000.0, np.nan]
})
print(df)
# name age salary
# 0 Alice 25.0 50000.0
# 1 None NaN 60000.0
# 2 Carol 30.0 NaNisna() and isnull()
isna() and isnull() are completely identical — both return a DataFrame or Series of the same shape filled with True wherever the value is missing and False elsewhere. Pandas provides both names purely for user preference. The result can be used directly as a boolean mask for filtering or further computation.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'A': [1, np.nan, 3],
'B': [np.nan, 2, np.nan]
})
print(df.isna())
# A B
# 0 False True
# 1 True False
# 2 False True
# Both are identical
print((df.isna() == df.isnull()).all().all()) # Truenotna() to Find Non-Missing Values
notna() (also aliased as notnull()) is the inverse of isna() — it returns True where values are present and False where they are missing. This is useful when you want to filter to rows that have a value in a critical column, such as requiring that a primary key or target variable is not NaN.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'id': [1, 2, 3],
'email': ['a@x.com', None, 'c@x.com']
})
# Keep only rows where email is present
with_email = df[df['email'].notna()]
print(with_email)
# id email
# 0 1 a@x.com
# 2 3 c@x.comCounting Missing Values per Column
Calling .isna().sum() on a DataFrame sums the True values (which equal 1) column-by-column, giving you the count of missing values per column. This is the single most useful first step in understanding the quality of a new dataset — it tells you which columns need attention.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'name': ['Alice', None, 'Carol', None],
'age': [25, np.nan, 30, 22],
'salary': [50000, 60000, np.nan, np.nan]
})
missing_counts = df.isna().sum()
print(missing_counts)
# name 2
# age 1
# salary 2
# dtype: int64Missing Percentage per Column
An absolute count of NaN values is less useful than the percentage of missing values because it scales with dataset size. Dividing isna().sum() by the total row count (or calling isna().mean()) gives the fraction missing, which you can multiply by 100 for a percentage. Columns with more than 30-50% missing often require a decision about whether to keep them at all.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'A': [1, np.nan, 3, np.nan, 5],
'B': [np.nan, 2, np.nan, 4, 5],
'C': [1, 2, 3, 4, 5]
})
missing_pct = (df.isna().mean() * 100).round(1)
print(missing_pct)
# A 40.0
# B 40.0
# C 0.0
# dtype: float64Missing Summary Table
A common EDA pattern is to build a missing value summary table that shows count, percentage, and dtype for each column in one view. This gives a complete picture of data quality before making any cleaning decisions. You can sort it to surface the most problematic columns first.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'col_a': [1, np.nan, 3],
'col_b': [np.nan, np.nan, 3],
'col_c': [1, 2, 3]
})
summary = pd.DataFrame({
'missing_count': df.isna().sum(),
'missing_pct': (df.isna().mean() * 100).round(1),
'dtype': df.dtypes
}).sort_values('missing_pct', ascending=False)
print(summary)
# missing_count missing_pct dtype
# col_b 2 66.7 float64
# col_a 1 33.3 float64
# col_c 0 0.0 float64Row-Level Missing Value Count
You can also count missing values per row by calling isna().sum(axis=1). This helps identify records that are mostly empty (e.g., incomplete survey responses) which you might want to flag or remove as a unit. A row with many missing values is fundamentally different from scattered column-level missingness.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'a': [1, np.nan, np.nan],
'b': [np.nan, 2, np.nan],
'c': [3, 4, np.nan]
})
# Count NaN per row
df['missing_count'] = df.isna().sum(axis=1)
print(df)
# a b c missing_count
# 0 1.0 NaN 3.0 1
# 1 NaN 2.0 4.0 1
# 2 NaN NaN NaN 3Filtering Rows with Any or All Missing
Use df[df.isna().any(axis=1)] to find rows that have at least one NaN value, or df[df.isna().all(axis=1)] to find rows where every value is NaN. These filters help you isolate problem records for inspection before deciding how to handle them.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'x': [1, np.nan, np.nan],
'y': [2, 3, np.nan],
'z': [4, 5, np.nan]
})
# Rows with at least one NaN
has_any_nan = df[df.isna().any(axis=1)]
print('Any NaN:\n', has_any_nan)
# Rows where all values are NaN
all_nan = df[df.isna().all(axis=1)]
print('All NaN:\n', all_nan)
# x y z
# 2 NaN NaN NaNChecking a Specific Column for NaN
For a quick sanity check on a single column, call df['col'].isna().sum() or use df['col'].isna().any() to get a single boolean (True if any NaN exists). These one-liners are useful inside data validation checks or logging statements in a pipeline.
import pandas as pd
import numpy as np
df = pd.DataFrame({'price': [10.0, np.nan, 30.0, np.nan, 50.0]})
print('NaN count in price:', df['price'].isna().sum()) # 2
print('Any NaN in price?', df['price'].isna().any()) # True
print('All present?', df['price'].notna().all()) # FalseVisualising Missing Values with a Heatmap
For datasets with many columns, a missing value heatmap is more informative than a table of numbers. You can create one easily with Seaborn: the darker a cell, the more missing data in that column-row combination. A popular third-party library called missingno provides dedicated missing-value visualisations.
import pandas as pd
import numpy as np
import seaborn as sns
import matplotlib.pyplot as plt
np.random.seed(0)
df = pd.DataFrame(
np.where(np.random.rand(20, 5) > 0.7, np.nan, np.random.randn(20, 5)),
columns=['A', 'B', 'C', 'D', 'E']
)
# Heatmap of missing values
sns.heatmap(df.isna(), cbar=False, yticklabels=False)
plt.title('Missing Value Pattern')
plt.tight_layout()
plt.savefig('missing_heatmap.png')
print('Saved missing_heatmap.png')info() for a Quick Missing Check
df.info() prints a concise summary that includes the non-null count for every column. This is the fastest way to spot columns with missing values in a new dataset: any column whose non-null count is less than the total row count has NaN values. It also shows dtype and memory usage.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'id': [1, 2, 3, 4, 5],
'age': [25, np.nan, 30, np.nan, 22],
'salary': [50000, 60000, np.nan, 70000, 80000]
})
df.info()
# <class 'pandas.core.frame.DataFrame'>
# RangeIndex: 5 entries, 0 to 4
# Data columns (total 3 columns):
# # Column Non-Null Count Dtype
# --- ------ -------------- -----
# 0 id 5 non-null int64
# 1 age 3 non-null float64
# 2 salary 4 non-null float64Quick Check
Test your understanding of detecting missing values in Pandas.
Lesson Recap
In this lesson you learned: isna() and isnull() are identical and return boolean masks of missing positions, isna().sum() counts NaN per column, and isna().mean()*100 gives the missing percentage. Use df.info() for a fast overview and isna().any(axis=1) to find rows with at least one NaN. Next up we tackle dropping missing values with dropna().
よくある質問
「欠損値の検出」レッスンは無料ですか?
はい。「欠損値の検出」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Pandas & NumPy Academyコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Pandas & NumPy Academyコースには全4レッスンが含まれています。
「欠損値の検出」で何を学びますか?
isna()、notna()、isnull()を使って、SeriesまたはDataFrame内のNaNの位置を見つけ、列ごとの欠損値を数えます。 ブラウザで直接実行するハンズオンコードでPandas & NumPy Academyを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。
Pandas & NumPy Academyを始めるのに経験は必要ですか?
事前経験は必要ありません。CoddyKitのPandas & NumPy Academyは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン1/4です。
「欠損値の検出」レッスンにはどのくらい時間がかかりますか?
ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。
このPandas & NumPy Academyレッスンでコードを書いて実行できますか?
はい。すべてのPandas & NumPy Academyレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。
このコースのすべてのレッスン
- 欠損値の検出
- 欠損値の削除
- 欠損値の補完
- 補間と高度な欠測値補完