0Pricing
Pandas & NumPy Academy · Lesson

Dataset Profiling Checklist

Systematically inspect shape, dtypes, missing counts, unique values, and basic statistics when encountering a new dataset.

Dataset Profiling Checklist is a free Pandas & NumPy Academy lesson on CoddyKit — lesson 1 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Pandas & NumPy Academy learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.

The First Five Minutes with a Dataset

Experienced data analysts follow a systematic profiling checklist whenever they encounter a new dataset. Rather than diving straight into analysis, they first answer a set of diagnostic questions: How large is the data? What types are the columns? How many values are missing? Are there obvious anomalies? This structured approach prevents hours of wasted work caused by misunderstood data types or hidden nulls corrupting calculations.

import pandas as pd

# Simulate loading a new dataset
df = pd.read_csv('https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv')

# Step 1: size
print('Shape:', df.shape)
print('Rows:', len(df))

Step 1: Shape, Columns, and dtypes

The first three checks are always shape (rows × columns), column names (do they match expectations?), and dtypes (are numeric columns actually numeric?). A common surprise is that numeric columns were loaded as object dtype because they contain text entries like 'unknown' or mixed comma-formatted numbers. Catching this at the profiling stage saves you from aggregations that silently return wrong results.

import pandas as pd
import seaborn as sns

df = sns.load_dataset('titanic')

print('Shape:', df.shape)
print('\nColumns:', df.columns.tolist())
print('\nData Types:')
print(df.dtypes)

Step 2: Missing Value Audit

Call df.isna().sum() to count missing values per column, then divide by len(df) to get the percentage. Columns with >50% missing are usually candidates for removal; 1-20% missing typically warrant imputation; 0% missing needs a sanity check — perfect completeness can itself be suspicious if real data rarely comes that clean. Sort the output descending to see the worst-affected columns first.

import pandas as pd
import seaborn as sns

df = sns.load_dataset('titanic')

# Missing value audit
missing = df.isna().sum()
missing_pct = (missing / len(df) * 100).round(1)
audit = pd.DataFrame({'count': missing, 'pct': missing_pct})
audit = audit[audit['count'] > 0].sort_values('pct', ascending=False)
print(audit)

Step 3: Basic Descriptive Statistics

df.describe() returns count, mean, std, min, quartiles, and max for every numeric column in one call. Always scan the min and max for impossible values (e.g. age = -3 or age = 999), the mean vs. median for skewness (large gap suggests outliers), and the count (if lower than total rows, there are nulls). Add include='all' to include statistics for object columns too (unique count, top, freq).

import pandas as pd
import seaborn as sns

df = sns.load_dataset('titanic')

# Describe numeric columns
print('Numeric statistics:')
print(df.describe().round(2))

print('\nObject column statistics:')
print(df.describe(include='object'))

Step 4: Unique Value Counts

For categorical columns, check the number of unique values with df[col].nunique() and inspect the most common values with value_counts(). Watch for columns that appear categorical but have hundreds of unique values (possibly a free-text field or an ID), and columns that should be boolean but have three values (True, False, and NaN). Also check for inconsistent capitalisation: 'Male' and 'male' treated as separate categories.

import pandas as pd
import seaborn as sns

df = sns.load_dataset('titanic')

categorical = df.select_dtypes(include='object').columns
for col in categorical:
    n = df[col].nunique()
    top_vals = df[col].value_counts().head(3).to_dict()
    print(f'{col}: {n} unique — top3: {top_vals}')

Step 5: Sample Rows with head and tail

df.head(10) and df.tail(10) show the first and last rows respectively. Examining both is important because datasets often have clean header rows and messy trailing rows from export artefacts (e.g. a 'Total' summary row appended to a CSV export). Also use df.sample(10) to see a random selection of rows, which avoids any bias toward the beginning of the dataset.

import pandas as pd
import seaborn as sns

df = sns.load_dataset('titanic')

print('First 3 rows:')
print(df.head(3))

print('\nLast 3 rows:')
print(df.tail(3))

print('\nRandom sample of 3:')
print(df.sample(3, random_state=42))

Step 6: Check for Duplicates

Call df.duplicated().sum() to count fully duplicated rows (where every column matches). A dataset claiming to contain unique customer records with even one duplicate is a red flag for data quality issues. Use df[df.duplicated(keep=False)] to inspect the actual duplicate rows. When a table should have unique keys (like a user ID), also check key-level uniqueness with df['id'].duplicated().sum().

import pandas as pd
import seaborn as sns

df = sns.load_dataset('titanic')

# Full-row duplicates
full_dups = df.duplicated().sum()
print(f'Fully duplicated rows: {full_dups}')

# Key-level duplicates (e.g. passenger class + name)
key_dups = df.duplicated(subset=['pclass', 'sex', 'age', 'fare']).sum()
print(f'Near-duplicate rows (pclass+sex+age+fare): {key_dups}')

Step 7: Date and Time Columns

If the dataset contains date or timestamp columns, check that Pandas loaded them as datetime64 (not object). Convert with pd.to_datetime(df['col']). Then inspect the date range: df['date'].min() and df['date'].max(). Look for timestamps far in the future (year 2099) or past (epoch zero: 1970-01-01) which indicate sentinel values used to represent missing dates.

import pandas as pd

# Synthetic example with a date column
df = pd.DataFrame({
    'order_id': [1, 2, 3, 4],
    'order_date': ['2024-01-10', '2024-02-05', '1970-01-01', '2024-12-31']
})

df['order_date'] = pd.to_datetime(df['order_date'])

print('Date range:', df['order_date'].min(), 'to', df['order_date'].max())
print('\nSuspect dates (before 2020):')
print(df[df['order_date'] < '2020-01-01'])

Automating the Checklist into a Function

Wrapping the profiling steps into a reusable function ensures the same checks run consistently on every dataset. The function should print a structured report showing shape, missing values, dtypes, duplicate count, and basic stats. You can extend it with visualisations or save it as an HTML report. Functions like this are a staple of professional data science teams where multiple analysts work on the same pipeline.

import pandas as pd

def profile(df, name='DataFrame'):
    print(f'=== Profile: {name} ===')
    print(f'Shape: {df.shape[0]} rows x {df.shape[1]} cols')
    print(f'Duplicates: {df.duplicated().sum()}')
    print(f'\nMissing values (top 5):')
    missing = (df.isna().sum() / len(df) * 100).sort_values(ascending=False)
    print(missing[missing > 0].head(5).round(1).to_string())
    print(f'\ndtypes:')
    print(df.dtypes.value_counts().to_string())
    print('=' * 35)

import seaborn as sns
df = sns.load_dataset('titanic')
profile(df, 'Titanic')

ydata-profiling for Automated EDA

The ydata-profiling library (formerly pandas-profiling) generates a comprehensive HTML report from a DataFrame with a single line: ProfileReport(df).to_file('report.html'). The report includes histograms, correlation matrices, missing value heatmaps, duplicate detection, and interaction plots. It is the fastest way to share a full dataset profile with a stakeholder who needs to understand the data without writing code.

# Install: pip install ydata-profiling
# from ydata_profiling import ProfileReport
# import pandas as pd
# import seaborn as sns

# df = sns.load_dataset('titanic')
# profile = ProfileReport(df, title='Titanic Profiling Report', explorative=True)
# profile.to_file('titanic_profile.html')

# The above generates a full HTML report automatically.
# Alternatively, use minimal mode for faster generation:
# profile = ProfileReport(df, minimal=True)
print('ydata-profiling generates HTML EDA reports automatically.')
print('Run: pip install ydata-profiling')

Profiling Checklist Summary

The complete profiling checklist for any new dataset contains seven steps: 1) shape and columns, 2) dtypes and suspicious type assignments, 3) missing value count and percentage per column, 4) descriptive statistics (min, max, mean, median), 5) unique values and value counts for categoricals, 6) duplicate row detection, and 7) date range validation. Completing this checklist before any analysis prevents silent errors that corrupt results downstream.

Quick Check

Test your understanding of the dataset profiling checklist from this lesson.

Lesson Recap

In this lesson you learned: the 7-step profiling checklist (shape, dtypes, missing values, stats, unique counts, duplicates, dates), wrapping checks into a reusable profile function, and using ydata-profiling for automated HTML reports. Next up we dive into univariate analysis — studying each column independently to spot outliers, skewness, and unusual distributions.

Frequently asked questions

Is the “Dataset Profiling Checklist” lesson free?

Yes — the full text of “Dataset Profiling Checklist” is free to read here on the web, and the Pandas & NumPy Academy course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Pandas & NumPy Academy course, upgrade to CoddyKit PRO.

What will I learn in “Dataset Profiling Checklist”?

Systematically inspect shape, dtypes, missing counts, unique values, and basic statistics when encountering a new dataset. You practise Pandas & NumPy Academy with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.

Do I need any experience to start Pandas & NumPy Academy?

No prior experience is required. Pandas & NumPy Academy on CoddyKit is structured for beginners through advanced learners; this is — lesson 1 of 4, so you can start here or from the beginning and move at your own pace.

How long does the “Dataset Profiling Checklist” lesson take?

Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.

Can I write and run code in this Pandas & NumPy Academy lesson?

Yes. Every Pandas & NumPy Academy lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.

All lessons in this course

  1. Dataset Profiling Checklist
  2. Univariate Analysis
  3. Bivariate and Correlation Analysis
  4. Summarising Findings in a Report
← Back to Pandas & NumPy Academy