0Pricing
Pandas & NumPy Academy · レッスン

データセットプロファイリングのチェックリスト

新しいデータセットを扱うときに、形状、dtypes、欠損数、一意な値、基本統計量を体系的に確認します。

「データセットプロファイリングのチェックリスト」はCoddyKit上の無料Pandas & NumPy Academyレッスンです。 これはレッスン1/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはPandas & NumPy Academy学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Pandas & NumPy Academyコースには全4レッスンが含まれています。

このレッスンの一部はまだ翻訳されておらず、英語で表示されています。

The First Five Minutes with a Dataset

Experienced data analysts follow a systematic profiling checklist whenever they encounter a new dataset. Rather than diving straight into analysis, they first answer a set of diagnostic questions: How large is the data? What types are the columns? How many values are missing? Are there obvious anomalies? This structured approach prevents hours of wasted work caused by misunderstood data types or hidden nulls corrupting calculations.

import pandas as pd

# Simulate loading a new dataset
df = pd.read_csv('https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv')

# Step 1: size
print('Shape:', df.shape)
print('Rows:', len(df))

Step 1: Shape, Columns, and dtypes

The first three checks are always shape (rows × columns), column names (do they match expectations?), and dtypes (are numeric columns actually numeric?). A common surprise is that numeric columns were loaded as object dtype because they contain text entries like 'unknown' or mixed comma-formatted numbers. Catching this at the profiling stage saves you from aggregations that silently return wrong results.

import pandas as pd
import seaborn as sns

df = sns.load_dataset('titanic')

print('Shape:', df.shape)
print('\nColumns:', df.columns.tolist())
print('\nData Types:')
print(df.dtypes)

Step 2: Missing Value Audit

Call df.isna().sum() to count missing values per column, then divide by len(df) to get the percentage. Columns with >50% missing are usually candidates for removal; 1-20% missing typically warrant imputation; 0% missing needs a sanity check — perfect completeness can itself be suspicious if real data rarely comes that clean. Sort the output descending to see the worst-affected columns first.

import pandas as pd
import seaborn as sns

df = sns.load_dataset('titanic')

# Missing value audit
missing = df.isna().sum()
missing_pct = (missing / len(df) * 100).round(1)
audit = pd.DataFrame({'count': missing, 'pct': missing_pct})
audit = audit[audit['count'] > 0].sort_values('pct', ascending=False)
print(audit)

Step 3: Basic Descriptive Statistics

df.describe() returns count, mean, std, min, quartiles, and max for every numeric column in one call. Always scan the min and max for impossible values (e.g. age = -3 or age = 999), the mean vs. median for skewness (large gap suggests outliers), and the count (if lower than total rows, there are nulls). Add include='all' to include statistics for object columns too (unique count, top, freq).

import pandas as pd
import seaborn as sns

df = sns.load_dataset('titanic')

# Describe numeric columns
print('Numeric statistics:')
print(df.describe().round(2))

print('\nObject column statistics:')
print(df.describe(include='object'))

Step 4: Unique Value Counts

For categorical columns, check the number of unique values with df[col].nunique() and inspect the most common values with value_counts(). Watch for columns that appear categorical but have hundreds of unique values (possibly a free-text field or an ID), and columns that should be boolean but have three values (True, False, and NaN). Also check for inconsistent capitalisation: 'Male' and 'male' treated as separate categories.

import pandas as pd
import seaborn as sns

df = sns.load_dataset('titanic')

categorical = df.select_dtypes(include='object').columns
for col in categorical:
    n = df[col].nunique()
    top_vals = df[col].value_counts().head(3).to_dict()
    print(f'{col}: {n} unique — top3: {top_vals}')

Step 5: Sample Rows with head and tail

df.head(10) and df.tail(10) show the first and last rows respectively. Examining both is important because datasets often have clean header rows and messy trailing rows from export artefacts (e.g. a 'Total' summary row appended to a CSV export). Also use df.sample(10) to see a random selection of rows, which avoids any bias toward the beginning of the dataset.

import pandas as pd
import seaborn as sns

df = sns.load_dataset('titanic')

print('First 3 rows:')
print(df.head(3))

print('\nLast 3 rows:')
print(df.tail(3))

print('\nRandom sample of 3:')
print(df.sample(3, random_state=42))

Step 6: Check for Duplicates

Call df.duplicated().sum() to count fully duplicated rows (where every column matches). A dataset claiming to contain unique customer records with even one duplicate is a red flag for data quality issues. Use df[df.duplicated(keep=False)] to inspect the actual duplicate rows. When a table should have unique keys (like a user ID), also check key-level uniqueness with df['id'].duplicated().sum().

import pandas as pd
import seaborn as sns

df = sns.load_dataset('titanic')

# Full-row duplicates
full_dups = df.duplicated().sum()
print(f'Fully duplicated rows: {full_dups}')

# Key-level duplicates (e.g. passenger class + name)
key_dups = df.duplicated(subset=['pclass', 'sex', 'age', 'fare']).sum()
print(f'Near-duplicate rows (pclass+sex+age+fare): {key_dups}')

Step 7: Date and Time Columns

If the dataset contains date or timestamp columns, check that Pandas loaded them as datetime64 (not object). Convert with pd.to_datetime(df['col']). Then inspect the date range: df['date'].min() and df['date'].max(). Look for timestamps far in the future (year 2099) or past (epoch zero: 1970-01-01) which indicate sentinel values used to represent missing dates.

import pandas as pd

# Synthetic example with a date column
df = pd.DataFrame({
    'order_id': [1, 2, 3, 4],
    'order_date': ['2024-01-10', '2024-02-05', '1970-01-01', '2024-12-31']
})

df['order_date'] = pd.to_datetime(df['order_date'])

print('Date range:', df['order_date'].min(), 'to', df['order_date'].max())
print('\nSuspect dates (before 2020):')
print(df[df['order_date'] < '2020-01-01'])

Automating the Checklist into a Function

Wrapping the profiling steps into a reusable function ensures the same checks run consistently on every dataset. The function should print a structured report showing shape, missing values, dtypes, duplicate count, and basic stats. You can extend it with visualisations or save it as an HTML report. Functions like this are a staple of professional data science teams where multiple analysts work on the same pipeline.

import pandas as pd

def profile(df, name='DataFrame'):
    print(f'=== Profile: {name} ===')
    print(f'Shape: {df.shape[0]} rows x {df.shape[1]} cols')
    print(f'Duplicates: {df.duplicated().sum()}')
    print(f'\nMissing values (top 5):')
    missing = (df.isna().sum() / len(df) * 100).sort_values(ascending=False)
    print(missing[missing > 0].head(5).round(1).to_string())
    print(f'\ndtypes:')
    print(df.dtypes.value_counts().to_string())
    print('=' * 35)

import seaborn as sns
df = sns.load_dataset('titanic')
profile(df, 'Titanic')

ydata-profiling for Automated EDA

The ydata-profiling library (formerly pandas-profiling) generates a comprehensive HTML report from a DataFrame with a single line: ProfileReport(df).to_file('report.html'). The report includes histograms, correlation matrices, missing value heatmaps, duplicate detection, and interaction plots. It is the fastest way to share a full dataset profile with a stakeholder who needs to understand the data without writing code.

# Install: pip install ydata-profiling
# from ydata_profiling import ProfileReport
# import pandas as pd
# import seaborn as sns

# df = sns.load_dataset('titanic')
# profile = ProfileReport(df, title='Titanic Profiling Report', explorative=True)
# profile.to_file('titanic_profile.html')

# The above generates a full HTML report automatically.
# Alternatively, use minimal mode for faster generation:
# profile = ProfileReport(df, minimal=True)
print('ydata-profiling generates HTML EDA reports automatically.')
print('Run: pip install ydata-profiling')

Profiling Checklist Summary

The complete profiling checklist for any new dataset contains seven steps: 1) shape and columns, 2) dtypes and suspicious type assignments, 3) missing value count and percentage per column, 4) descriptive statistics (min, max, mean, median), 5) unique values and value counts for categoricals, 6) duplicate row detection, and 7) date range validation. Completing this checklist before any analysis prevents silent errors that corrupt results downstream.

Quick Check

Test your understanding of the dataset profiling checklist from this lesson.

Lesson Recap

In this lesson you learned: the 7-step profiling checklist (shape, dtypes, missing values, stats, unique counts, duplicates, dates), wrapping checks into a reusable profile function, and using ydata-profiling for automated HTML reports. Next up we dive into univariate analysis — studying each column independently to spot outliers, skewness, and unusual distributions.

よくある質問

「データセットプロファイリングのチェックリスト」レッスンは無料ですか?

はい。「データセットプロファイリングのチェックリスト」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Pandas & NumPy Academyコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Pandas & NumPy Academyコースには全4レッスンが含まれています。

「データセットプロファイリングのチェックリスト」で何を学びますか?

新しいデータセットを扱うときに、形状、dtypes、欠損数、一意な値、基本統計量を体系的に確認します。 ブラウザで直接実行するハンズオンコードでPandas & NumPy Academyを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。

Pandas & NumPy Academyを始めるのに経験は必要ですか?

事前経験は必要ありません。CoddyKitのPandas & NumPy Academyは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン1/4です。

「データセットプロファイリングのチェックリスト」レッスンにはどのくらい時間がかかりますか?

ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。

このPandas & NumPy Academyレッスンでコードを書いて実行できますか?

はい。すべてのPandas & NumPy Academyレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。

このコースのすべてのレッスン

  1. データセットプロファイリングのチェックリスト
  2. 一変量解析
  3. 二変量解析と相関分析
  4. レポートでの分析結果の要約
← Pandas & NumPy Academyに戻る