Lista de comprobación para perfilar un dataset
Inspeccione sistemáticamente la forma, los tipos de datos, los recuentos de valores faltantes, los valores únicos y las estadísticas básicas al trabajar con un dataset nuevo.
Lista de comprobación para perfilar un dataset es una lección gratuita de Pandas & NumPy Academy en CoddyKit. Esta es la lección 1 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Pandas & NumPy Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Pandas & NumPy Academy incluye 4 lecciones en total.
Partes de esta lección aún no han sido traducidas y se muestran en inglés.
The First Five Minutes with a Dataset
Experienced data analysts follow a systematic profiling checklist whenever they encounter a new dataset. Rather than diving straight into analysis, they first answer a set of diagnostic questions: How large is the data? What types are the columns? How many values are missing? Are there obvious anomalies? This structured approach prevents hours of wasted work caused by misunderstood data types or hidden nulls corrupting calculations.
import pandas as pd
# Simulate loading a new dataset
df = pd.read_csv('https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv')
# Step 1: size
print('Shape:', df.shape)
print('Rows:', len(df))Step 1: Shape, Columns, and dtypes
The first three checks are always shape (rows × columns), column names (do they match expectations?), and dtypes (are numeric columns actually numeric?). A common surprise is that numeric columns were loaded as object dtype because they contain text entries like 'unknown' or mixed comma-formatted numbers. Catching this at the profiling stage saves you from aggregations that silently return wrong results.
import pandas as pd
import seaborn as sns
df = sns.load_dataset('titanic')
print('Shape:', df.shape)
print('\nColumns:', df.columns.tolist())
print('\nData Types:')
print(df.dtypes)Step 2: Missing Value Audit
Call df.isna().sum() to count missing values per column, then divide by len(df) to get the percentage. Columns with >50% missing are usually candidates for removal; 1-20% missing typically warrant imputation; 0% missing needs a sanity check — perfect completeness can itself be suspicious if real data rarely comes that clean. Sort the output descending to see the worst-affected columns first.
import pandas as pd
import seaborn as sns
df = sns.load_dataset('titanic')
# Missing value audit
missing = df.isna().sum()
missing_pct = (missing / len(df) * 100).round(1)
audit = pd.DataFrame({'count': missing, 'pct': missing_pct})
audit = audit[audit['count'] > 0].sort_values('pct', ascending=False)
print(audit)Step 3: Basic Descriptive Statistics
df.describe() returns count, mean, std, min, quartiles, and max for every numeric column in one call. Always scan the min and max for impossible values (e.g. age = -3 or age = 999), the mean vs. median for skewness (large gap suggests outliers), and the count (if lower than total rows, there are nulls). Add include='all' to include statistics for object columns too (unique count, top, freq).
import pandas as pd
import seaborn as sns
df = sns.load_dataset('titanic')
# Describe numeric columns
print('Numeric statistics:')
print(df.describe().round(2))
print('\nObject column statistics:')
print(df.describe(include='object'))Step 4: Unique Value Counts
For categorical columns, check the number of unique values with df[col].nunique() and inspect the most common values with value_counts(). Watch for columns that appear categorical but have hundreds of unique values (possibly a free-text field or an ID), and columns that should be boolean but have three values (True, False, and NaN). Also check for inconsistent capitalisation: 'Male' and 'male' treated as separate categories.
import pandas as pd
import seaborn as sns
df = sns.load_dataset('titanic')
categorical = df.select_dtypes(include='object').columns
for col in categorical:
n = df[col].nunique()
top_vals = df[col].value_counts().head(3).to_dict()
print(f'{col}: {n} unique — top3: {top_vals}')Step 5: Sample Rows with head and tail
df.head(10) and df.tail(10) show the first and last rows respectively. Examining both is important because datasets often have clean header rows and messy trailing rows from export artefacts (e.g. a 'Total' summary row appended to a CSV export). Also use df.sample(10) to see a random selection of rows, which avoids any bias toward the beginning of the dataset.
import pandas as pd
import seaborn as sns
df = sns.load_dataset('titanic')
print('First 3 rows:')
print(df.head(3))
print('\nLast 3 rows:')
print(df.tail(3))
print('\nRandom sample of 3:')
print(df.sample(3, random_state=42))Step 6: Check for Duplicates
Call df.duplicated().sum() to count fully duplicated rows (where every column matches). A dataset claiming to contain unique customer records with even one duplicate is a red flag for data quality issues. Use df[df.duplicated(keep=False)] to inspect the actual duplicate rows. When a table should have unique keys (like a user ID), also check key-level uniqueness with df['id'].duplicated().sum().
import pandas as pd
import seaborn as sns
df = sns.load_dataset('titanic')
# Full-row duplicates
full_dups = df.duplicated().sum()
print(f'Fully duplicated rows: {full_dups}')
# Key-level duplicates (e.g. passenger class + name)
key_dups = df.duplicated(subset=['pclass', 'sex', 'age', 'fare']).sum()
print(f'Near-duplicate rows (pclass+sex+age+fare): {key_dups}')Step 7: Date and Time Columns
If the dataset contains date or timestamp columns, check that Pandas loaded them as datetime64 (not object). Convert with pd.to_datetime(df['col']). Then inspect the date range: df['date'].min() and df['date'].max(). Look for timestamps far in the future (year 2099) or past (epoch zero: 1970-01-01) which indicate sentinel values used to represent missing dates.
import pandas as pd
# Synthetic example with a date column
df = pd.DataFrame({
'order_id': [1, 2, 3, 4],
'order_date': ['2024-01-10', '2024-02-05', '1970-01-01', '2024-12-31']
})
df['order_date'] = pd.to_datetime(df['order_date'])
print('Date range:', df['order_date'].min(), 'to', df['order_date'].max())
print('\nSuspect dates (before 2020):')
print(df[df['order_date'] < '2020-01-01'])Automating the Checklist into a Function
Wrapping the profiling steps into a reusable function ensures the same checks run consistently on every dataset. The function should print a structured report showing shape, missing values, dtypes, duplicate count, and basic stats. You can extend it with visualisations or save it as an HTML report. Functions like this are a staple of professional data science teams where multiple analysts work on the same pipeline.
import pandas as pd
def profile(df, name='DataFrame'):
print(f'=== Profile: {name} ===')
print(f'Shape: {df.shape[0]} rows x {df.shape[1]} cols')
print(f'Duplicates: {df.duplicated().sum()}')
print(f'\nMissing values (top 5):')
missing = (df.isna().sum() / len(df) * 100).sort_values(ascending=False)
print(missing[missing > 0].head(5).round(1).to_string())
print(f'\ndtypes:')
print(df.dtypes.value_counts().to_string())
print('=' * 35)
import seaborn as sns
df = sns.load_dataset('titanic')
profile(df, 'Titanic')ydata-profiling for Automated EDA
The ydata-profiling library (formerly pandas-profiling) generates a comprehensive HTML report from a DataFrame with a single line: ProfileReport(df).to_file('report.html'). The report includes histograms, correlation matrices, missing value heatmaps, duplicate detection, and interaction plots. It is the fastest way to share a full dataset profile with a stakeholder who needs to understand the data without writing code.
# Install: pip install ydata-profiling
# from ydata_profiling import ProfileReport
# import pandas as pd
# import seaborn as sns
# df = sns.load_dataset('titanic')
# profile = ProfileReport(df, title='Titanic Profiling Report', explorative=True)
# profile.to_file('titanic_profile.html')
# The above generates a full HTML report automatically.
# Alternatively, use minimal mode for faster generation:
# profile = ProfileReport(df, minimal=True)
print('ydata-profiling generates HTML EDA reports automatically.')
print('Run: pip install ydata-profiling')Profiling Checklist Summary
The complete profiling checklist for any new dataset contains seven steps: 1) shape and columns, 2) dtypes and suspicious type assignments, 3) missing value count and percentage per column, 4) descriptive statistics (min, max, mean, median), 5) unique values and value counts for categoricals, 6) duplicate row detection, and 7) date range validation. Completing this checklist before any analysis prevents silent errors that corrupt results downstream.
Quick Check
Test your understanding of the dataset profiling checklist from this lesson.
Lesson Recap
In this lesson you learned: the 7-step profiling checklist (shape, dtypes, missing values, stats, unique counts, duplicates, dates), wrapping checks into a reusable profile function, and using ydata-profiling for automated HTML reports. Next up we dive into univariate analysis — studying each column independently to spot outliers, skewness, and unusual distributions.
Preguntas frecuentes
¿La lección «Lista de comprobación para perfilar un dataset» es gratis?
Sí — el texto completo de «Lista de comprobación para perfilar un dataset» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Pandas & NumPy Academy, actualiza a CoddyKit PRO. El curso de Pandas & NumPy Academy incluye 4 lecciones en total.
¿Qué aprenderé en «Lista de comprobación para perfilar un dataset»?
Inspeccione sistemáticamente la forma, los tipos de datos, los recuentos de valores faltantes, los valores únicos y las estadísticas básicas al trabajar con un dataset nuevo. Practicas Pandas & NumPy Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.
¿Necesito experiencia previa para empezar Pandas & NumPy Academy?
No se requiere experiencia previa. Pandas & NumPy Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 1 de 4.
¿Cuánto tiempo toma la lección «Lista de comprobación para perfilar un dataset»?
La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.
¿Puedo escribir y ejecutar código en esta lección de Pandas & NumPy Academy?
Sí. Cada lección de Pandas & NumPy Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.
Todas las lecciones de este curso
- Lista de comprobación para perfilar un dataset
- Análisis univariante
- Análisis bivariante y de correlación
- Resumir conclusiones en un informe