数据集分析清单
遇到新数据集时,系统检查其形状、dtypes、缺失值计数、唯一值和基本统计量。
数据集分析清单 是 CoddyKit 上的免费 Pandas & NumPy Academy 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Pandas & NumPy Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Pandas & NumPy Academy 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
The First Five Minutes with a Dataset
Experienced data analysts follow a systematic profiling checklist whenever they encounter a new dataset. Rather than diving straight into analysis, they first answer a set of diagnostic questions: How large is the data? What types are the columns? How many values are missing? Are there obvious anomalies? This structured approach prevents hours of wasted work caused by misunderstood data types or hidden nulls corrupting calculations.
import pandas as pd
# Simulate loading a new dataset
df = pd.read_csv('https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv')
# Step 1: size
print('Shape:', df.shape)
print('Rows:', len(df))Step 1: Shape, Columns, and dtypes
The first three checks are always shape (rows × columns), column names (do they match expectations?), and dtypes (are numeric columns actually numeric?). A common surprise is that numeric columns were loaded as object dtype because they contain text entries like 'unknown' or mixed comma-formatted numbers. Catching this at the profiling stage saves you from aggregations that silently return wrong results.
import pandas as pd
import seaborn as sns
df = sns.load_dataset('titanic')
print('Shape:', df.shape)
print('\nColumns:', df.columns.tolist())
print('\nData Types:')
print(df.dtypes)Step 2: Missing Value Audit
Call df.isna().sum() to count missing values per column, then divide by len(df) to get the percentage. Columns with >50% missing are usually candidates for removal; 1-20% missing typically warrant imputation; 0% missing needs a sanity check — perfect completeness can itself be suspicious if real data rarely comes that clean. Sort the output descending to see the worst-affected columns first.
import pandas as pd
import seaborn as sns
df = sns.load_dataset('titanic')
# Missing value audit
missing = df.isna().sum()
missing_pct = (missing / len(df) * 100).round(1)
audit = pd.DataFrame({'count': missing, 'pct': missing_pct})
audit = audit[audit['count'] > 0].sort_values('pct', ascending=False)
print(audit)Step 3: Basic Descriptive Statistics
df.describe() returns count, mean, std, min, quartiles, and max for every numeric column in one call. Always scan the min and max for impossible values (e.g. age = -3 or age = 999), the mean vs. median for skewness (large gap suggests outliers), and the count (if lower than total rows, there are nulls). Add include='all' to include statistics for object columns too (unique count, top, freq).
import pandas as pd
import seaborn as sns
df = sns.load_dataset('titanic')
# Describe numeric columns
print('Numeric statistics:')
print(df.describe().round(2))
print('\nObject column statistics:')
print(df.describe(include='object'))Step 4: Unique Value Counts
For categorical columns, check the number of unique values with df[col].nunique() and inspect the most common values with value_counts(). Watch for columns that appear categorical but have hundreds of unique values (possibly a free-text field or an ID), and columns that should be boolean but have three values (True, False, and NaN). Also check for inconsistent capitalisation: 'Male' and 'male' treated as separate categories.
import pandas as pd
import seaborn as sns
df = sns.load_dataset('titanic')
categorical = df.select_dtypes(include='object').columns
for col in categorical:
n = df[col].nunique()
top_vals = df[col].value_counts().head(3).to_dict()
print(f'{col}: {n} unique — top3: {top_vals}')Step 5: Sample Rows with head and tail
df.head(10) and df.tail(10) show the first and last rows respectively. Examining both is important because datasets often have clean header rows and messy trailing rows from export artefacts (e.g. a 'Total' summary row appended to a CSV export). Also use df.sample(10) to see a random selection of rows, which avoids any bias toward the beginning of the dataset.
import pandas as pd
import seaborn as sns
df = sns.load_dataset('titanic')
print('First 3 rows:')
print(df.head(3))
print('\nLast 3 rows:')
print(df.tail(3))
print('\nRandom sample of 3:')
print(df.sample(3, random_state=42))Step 6: Check for Duplicates
Call df.duplicated().sum() to count fully duplicated rows (where every column matches). A dataset claiming to contain unique customer records with even one duplicate is a red flag for data quality issues. Use df[df.duplicated(keep=False)] to inspect the actual duplicate rows. When a table should have unique keys (like a user ID), also check key-level uniqueness with df['id'].duplicated().sum().
import pandas as pd
import seaborn as sns
df = sns.load_dataset('titanic')
# Full-row duplicates
full_dups = df.duplicated().sum()
print(f'Fully duplicated rows: {full_dups}')
# Key-level duplicates (e.g. passenger class + name)
key_dups = df.duplicated(subset=['pclass', 'sex', 'age', 'fare']).sum()
print(f'Near-duplicate rows (pclass+sex+age+fare): {key_dups}')Step 7: Date and Time Columns
If the dataset contains date or timestamp columns, check that Pandas loaded them as datetime64 (not object). Convert with pd.to_datetime(df['col']). Then inspect the date range: df['date'].min() and df['date'].max(). Look for timestamps far in the future (year 2099) or past (epoch zero: 1970-01-01) which indicate sentinel values used to represent missing dates.
import pandas as pd
# Synthetic example with a date column
df = pd.DataFrame({
'order_id': [1, 2, 3, 4],
'order_date': ['2024-01-10', '2024-02-05', '1970-01-01', '2024-12-31']
})
df['order_date'] = pd.to_datetime(df['order_date'])
print('Date range:', df['order_date'].min(), 'to', df['order_date'].max())
print('\nSuspect dates (before 2020):')
print(df[df['order_date'] < '2020-01-01'])Automating the Checklist into a Function
Wrapping the profiling steps into a reusable function ensures the same checks run consistently on every dataset. The function should print a structured report showing shape, missing values, dtypes, duplicate count, and basic stats. You can extend it with visualisations or save it as an HTML report. Functions like this are a staple of professional data science teams where multiple analysts work on the same pipeline.
import pandas as pd
def profile(df, name='DataFrame'):
print(f'=== Profile: {name} ===')
print(f'Shape: {df.shape[0]} rows x {df.shape[1]} cols')
print(f'Duplicates: {df.duplicated().sum()}')
print(f'\nMissing values (top 5):')
missing = (df.isna().sum() / len(df) * 100).sort_values(ascending=False)
print(missing[missing > 0].head(5).round(1).to_string())
print(f'\ndtypes:')
print(df.dtypes.value_counts().to_string())
print('=' * 35)
import seaborn as sns
df = sns.load_dataset('titanic')
profile(df, 'Titanic')ydata-profiling for Automated EDA
The ydata-profiling library (formerly pandas-profiling) generates a comprehensive HTML report from a DataFrame with a single line: ProfileReport(df).to_file('report.html'). The report includes histograms, correlation matrices, missing value heatmaps, duplicate detection, and interaction plots. It is the fastest way to share a full dataset profile with a stakeholder who needs to understand the data without writing code.
# Install: pip install ydata-profiling
# from ydata_profiling import ProfileReport
# import pandas as pd
# import seaborn as sns
# df = sns.load_dataset('titanic')
# profile = ProfileReport(df, title='Titanic Profiling Report', explorative=True)
# profile.to_file('titanic_profile.html')
# The above generates a full HTML report automatically.
# Alternatively, use minimal mode for faster generation:
# profile = ProfileReport(df, minimal=True)
print('ydata-profiling generates HTML EDA reports automatically.')
print('Run: pip install ydata-profiling')Profiling Checklist Summary
The complete profiling checklist for any new dataset contains seven steps: 1) shape and columns, 2) dtypes and suspicious type assignments, 3) missing value count and percentage per column, 4) descriptive statistics (min, max, mean, median), 5) unique values and value counts for categoricals, 6) duplicate row detection, and 7) date range validation. Completing this checklist before any analysis prevents silent errors that corrupt results downstream.
Quick Check
Test your understanding of the dataset profiling checklist from this lesson.
Lesson Recap
In this lesson you learned: the 7-step profiling checklist (shape, dtypes, missing values, stats, unique counts, duplicates, dates), wrapping checks into a reusable profile function, and using ydata-profiling for automated HTML reports. Next up we dive into univariate analysis — studying each column independently to spot outliers, skewness, and unusual distributions.
常见问题解答
「数据集分析清单」课时是免费的吗?
是的 — 「数据集分析清单」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Pandas & NumPy Academy 课程的其余内容,请升级到 CoddyKit PRO。 Pandas & NumPy Academy 课程共包含 4 节课。
「数据集分析清单」这节课中我会学到什么?
遇到新数据集时,系统检查其形状、dtypes、缺失值计数、唯一值和基本统计量。 你通过在浏览器中直接运行的动手代码来练习 Pandas & NumPy Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Pandas & NumPy Academy 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Pandas & NumPy Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。
「数据集分析清单」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Pandas & NumPy Academy 课中编写并运行代码吗?
能。每节 Pandas & NumPy Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。